Whole-book editorial benchmark · V0.1

A book begins as fragments. Fred sees where they converge.

Fred builds a whole-book view from dispersed manuscript evidence. This benchmark tests whether that view produces more useful editorial decisions than a sophisticated Claude Opus 5 workflow.

80%Fred

20%Claude Opus 5

Highest-priority whole-book findings the writer would act onActionable = adopt or modify after a blind review.
Fred builds whole-book understanding through successive convergenceThirty-six manuscript fragments resolve into nine grouped representations of people, places and concepts, then into three higher-order editorial views of throughline, themes and tensions.PEOPLEPLACESCONCEPTSTHROUGHLINETHEMESTENSIONSFRED · WHOLE-BOOK STATEMANUSCRIPT FRAGMENTSMANUSCRIPT PROGRESSIONACCUMULATING WHOLE-BOOK CONTEXTThirty-six fragments resolve into nine middle-layer clusters, then three higher-order editorial views.

Results

What the benchmark found

The study keeps judgment, breadth, and deliverable editorial yield separate. It does not collapse performance into a single score.

Priority judgment

80% Fred · 20% Claude Opus 5

The share of each system's highest-priority whole-book findings the writer would adopt or modify after a blind review.

Editorial breadth

772 Fred · 201 Claude Opus 5

Editorial contributions returned across the manuscript. This measures how much work came back ready to inspect, not whether every contribution was correct.

Editorial yield

5.4× more professionally deliverable suggestions

Fred returned 161 professionally deliverable suggestions, versus 30 from Claude Opus 5. This measures the usable editorial work returned, rather than rewarding a shorter list for being easier to keep clean.

AI-detector result

99% Human before · 99% Human after

On a human-written book-length manuscript, accepting 73 executable Fred edits left Pangram 4.0's whole-book classification unchanged. A complete Claude Opus 5 rewrite of a frozen passage registered Mixed, confirming detector sensitivity.

Method

Fred and Claude Opus 5 received the same whole-book editorial assignment. Their returns were reviewed without system labels. Judgment, breadth, and professionally deliverable output were scored separately.

WordsmithWhole-book editorial benchmark · V0.1
The Wordsmith Whole-Book Editorial Benchmark, V0.1 | Wordsmith