The Evidence Benchmark
We planted the defects. Here's exactly what the audit caught — and proof it won't invent findings.
Most manuscript-analysis tools ask you to trust them. We'd rather show you. We wrote a short manuscript, deliberately planted a documented set of defects in it, ran the audit, and published the whole scorecard — the catches with their quoted evidence, the honest misses, and a test where we fed the system a fabricated finding to see if it would repeat it.
Everything here is reproducible. The manuscript is ours, written for this test, so we can publish every line. The answer key was written before the run.
Why this test, and why it's built this way
The problem with the category is hallucination. Independent reviews have documented AI manuscript tools inventing findings — attributing an action to the wrong character, citing a contradiction that isn't in the book. A report you have to fact-check is worse than no report. So the thing worth proving isn't "our AI is smart" — it's "every finding we hand you is anchored to a real line, and our system mechanically refuses to pass along a quote it can't find in your text." That's what this benchmark measures.
The hallucination test — the one that matters most
We took the audit's real output and added a plausible-sounding fabricated finding: a STRONG continuity flag claiming “Nella confessed that she had drowned her grandmother herself, beneath the harbor light” — a sentence that appears nowhere in the manuscript. Then we ran the evidence-anchor gate over the whole set.
“Nella confessed that she had drowned her grandmother herself, beneath the harbor light”
The gate string-matched the quote against the manuscript, found nothing, and demoted it out of the report — the only finding it demoted. All seven genuine findings, which quote real lines, were untouched. This is the mechanism that lets us make the guarantee below.
Real planted defects the audit caught — with the evidence it quoted
Planted: the day-counter runs backward — Chapter 4 is dated Day 6, Chapter 6 opens on Day 4.
Audit said: “Ch 4 day-counter reaches Day 6 but Ch 6 starts at Day 4, an apparent backward jump.”
Planted: the same 12-word sentence appears twice in Chapter 2.
“the tide went out and left the flats shining like hammered pewter”
Audit said: “12-token sequence repeats 2 times within Chapter 2.” — evidence verified against the text.
Found: the phrase “the lamp room” recurs three times across three chapters.
Audit said: “Phrase ‘the lamp room’ recurs 3 times across 3 chapters — possible motif or redundancy.” It flags, and tells you it might be intentional, rather than pretending to be certain.
And a repetition it correctly did not flag
We also planted an intentional refrain — “salt got into everything on the Reach, even the ledgers, even the dead” — and declared it on the intake form as a deliberate motif. The audit left it alone. Distinguishing an author's intentional echo from an accidental one is the difference between a useful redundancy report and a noisy one; a word-counting tool can't tell them apart, and this one is built to.
The honest part: the full planted-defect ledger
Here is every defect we planted, and exactly how this run scored on it — including what this deterministic layer missed and what belongs to other parts of the product. We publish the misses because a benchmark that only shows wins isn't a benchmark; it's an ad.
| Planted defect | Where | Result |
|---|---|---|
| Timeline runs backward (day-counter) | Ch 4→6 | CAUGHT — with evidence |
| Duplicated sentence within a chapter | Ch 2 | CAUGHT — with evidence |
| Phrase repeated across chapters | Ch 6–8 | CAUGHT — flagged for your call |
| Intentional refrain (declared on intake) | Ch 1, 7 | CORRECTLY IGNORED |
| Stock descriptor reused 3× (“the color of weak tea”) | Ch 1, 4, 8 | MISSED by this layer |
| Tense slip (present verb in past narration) | Ch 4 | MISSED here; the pass over-flagged 5 other lines it hands to the human/judge step |
| Character fact contradiction (Marren: never left ↔ war years in Lisbon) | Ch 1 ↔ 5 | Entity-state + judge layer — not exercised in this deterministic-only run |
| World fact contradiction (lighthouse: dark 30 years ↔ burned nightly) | Ch 1 ↔ 5 | Entity-state + judge layer |
| Object-state contradiction (ledger: sealed in evidence locker ↔ on her table) | Ch 2 ↔ 6 | Entity-state + judge layer |
| Co-presence impossibility (Halloran: left on last boat ↔ at the pier next morning) | Ch 4 ↔ 5 | Entity-state + judge layer |
| Undelivered promise (“the second key before the tides turned”) | Ch 2 | Promise-tracking + judge layer |
| Beat re-landed in different words | Ch 5 | Redundancy product (out of scope here) |
| Narrator over-explaining the theme (an AI tell) | Ch 8 | Dev-Edit product (out of scope here) |
Two honest notes. First, this run is the mechanical layer only — the entity-state contradiction detections (the character, world, object, co-presence and promise items above) are produced by the audit's Ledger-and-judge layer, which we didn't exercise in this reproducible-by-anyone run; we'll publish a second benchmark that does. Second, the deterministic tense pass is deliberately over-inclusive: it raised five candidates and left the decision to the review step, which is why it appears here as a limitation rather than a clean catch. We would rather tell you that than hide it.
This is why we can guarantee it. Because every finding is anchored to a line we can point to, we put money on it: if a top-severity finding in your report is ever shown to be factually wrong — the quote isn't in your book, or it misreads the passage — that report is free. The benchmark above is the mechanism that makes that promise safe to offer.