Priority judgment
80% Fred · 20% Claude Opus 5
The share of each system's highest-priority whole-book findings the writer would adopt or modify after a blind review.
Whole-book editorial benchmark · V0.1
Fred builds a whole-book view from dispersed manuscript evidence. This benchmark tests whether that view produces more useful editorial decisions than a sophisticated Claude Opus 5 workflow.
80%Fred
20%Claude Opus 5
Highest-priority whole-book findings the writer would act onActionable = adopt or modify after a blind review.Results
The study keeps judgment, breadth, and deliverable editorial yield separate. It does not collapse performance into a single score.
80% Fred · 20% Claude Opus 5
The share of each system's highest-priority whole-book findings the writer would adopt or modify after a blind review.
772 Fred · 201 Claude Opus 5
Editorial contributions returned across the manuscript. This measures how much work came back ready to inspect, not whether every contribution was correct.
5.4× more professionally deliverable suggestions
Fred returned 161 professionally deliverable suggestions, versus 30 from Claude Opus 5. This measures the usable editorial work returned, rather than rewarding a shorter list for being easier to keep clean.
99% Human before · 99% Human after
On a human-written book-length manuscript, accepting 73 executable Fred edits left Pangram 4.0's whole-book classification unchanged. A complete Claude Opus 5 rewrite of a frozen passage registered Mixed, confirming detector sensitivity.
Method
Fred and Claude Opus 5 received the same whole-book editorial assignment. Their returns were reviewed without system labels. Judgment, breadth, and professionally deliverable output were scored separately.