Threadright

The Evidence Benchmark

We planted the defects. Here's exactly what the audit caught — and proof it won't invent findings.

Most manuscript-analysis tools ask you to trust them. We'd rather show you. We wrote a short manuscript, deliberately planted a documented set of defects in it, ran the audit, and published the whole scorecard — the catches with their quoted evidence, the honest misses, and a test where we fed the system a fabricated finding to see if it would repeat it.

Everything here is reproducible. The manuscript is ours, written for this test, so we can publish every line. The answer key was written before the run.

7 / 7
findings quoted a real line from the manuscript, mechanically verified — zero cite text that isn't there
1 / 1
fabricated finding we injected was caught and rejected — the sole demotion, no real finding touched
100%
of delivered findings carry a chapter-and-line citation you can check in seconds

Why this test, and why it's built this way

The problem with the category is hallucination. Independent reviews have documented AI manuscript tools inventing findings — attributing an action to the wrong character, citing a contradiction that isn't in the book. A report you have to fact-check is worse than no report. So the thing worth proving isn't "our AI is smart" — it's "every finding we hand you is anchored to a real line, and our system mechanically refuses to pass along a quote it can't find in your text." That's what this benchmark measures.

Method (so you can repeat it). The manuscript is “The Salt Ledger,” a 9-chapter, ~1,460-word literary mystery we wrote for this test. Before running anything, we recorded a ground-truth ledger of 12 planted defects. We then ran the read's deterministic layer — the part that extracts state and repetition mechanically and verifies every quoted excerpt against the source text — and scored the output against the answer key. Finally we injected one fabricated finding, citing a sentence that does not appear in the manuscript, to confirm the evidence gate rejects it. This run exercises the mechanical, fully-reproducible layer; it does not stand in for the human editorial review every paid report also receives.

The hallucination test — the one that matters most

We took the audit's real output and added a plausible-sounding fabricated finding: a STRONG continuity flag claiming “Nella confessed that she had drowned her grandmother herself, beneath the harbor light” — a sentence that appears nowhere in the manuscript. Then we ran the evidence-anchor gate over the whole set.

Injected fabrication claimed Chapter 5 rejected → downgraded to OBSERVATION
“Nella confessed that she had drowned her grandmother herself, beneath the harbor light”

The gate string-matched the quote against the manuscript, found nothing, and demoted it out of the report — the only finding it demoted. All seven genuine findings, which quote real lines, were untouched. This is the mechanism that lets us make the guarantee below.

Real planted defects the audit caught — with the evidence it quoted

Timeline regression Chapters 4 → 6 ✓ caught

Planted: the day-counter runs backward — Chapter 4 is dated Day 6, Chapter 6 opens on Day 4.

Audit said: “Ch 4 day-counter reaches Day 6 but Ch 6 starts at Day 4, an apparent backward jump.”

Duplicated sentence (within a chapter) Chapter 2 ✓ caught

Planted: the same 12-word sentence appears twice in Chapter 2.

“the tide went out and left the flats shining like hammered pewter”

Audit said: “12-token sequence repeats 2 times within Chapter 2.” — evidence verified against the text.

Repeated phrase (across chapters) Chapters 6, 7, 8 ✓ caught

Found: the phrase “the lamp room” recurs three times across three chapters.

Audit said: “Phrase ‘the lamp room’ recurs 3 times across 3 chapters — possible motif or redundancy.” It flags, and tells you it might be intentional, rather than pretending to be certain.

And a repetition it correctly did not flag

We also planted an intentional refrain — “salt got into everything on the Reach, even the ledgers, even the dead” — and declared it on the intake form as a deliberate motif. The audit left it alone. Distinguishing an author's intentional echo from an accidental one is the difference between a useful redundancy report and a noisy one; a word-counting tool can't tell them apart, and this one is built to.


The honest part: the full planted-defect ledger

Here is every defect we planted, and exactly how this run scored on it — including what this deterministic layer missed and what belongs to other parts of the product. We publish the misses because a benchmark that only shows wins isn't a benchmark; it's an ad.

Planted defectWhereResult
Timeline runs backward (day-counter)Ch 4→6CAUGHT — with evidence
Duplicated sentence within a chapterCh 2CAUGHT — with evidence
Phrase repeated across chaptersCh 6–8CAUGHT — flagged for your call
Intentional refrain (declared on intake)Ch 1, 7CORRECTLY IGNORED
Stock descriptor reused 3× (“the color of weak tea”)Ch 1, 4, 8MISSED by this layer
Tense slip (present verb in past narration)Ch 4MISSED here; the pass over-flagged 5 other lines it hands to the human/judge step
Character fact contradiction (Marren: never left ↔ war years in Lisbon)Ch 1 ↔ 5Entity-state + judge layer — not exercised in this deterministic-only run
World fact contradiction (lighthouse: dark 30 years ↔ burned nightly)Ch 1 ↔ 5Entity-state + judge layer
Object-state contradiction (ledger: sealed in evidence locker ↔ on her table)Ch 2 ↔ 6Entity-state + judge layer
Co-presence impossibility (Halloran: left on last boat ↔ at the pier next morning)Ch 4 ↔ 5Entity-state + judge layer
Undelivered promise (“the second key before the tides turned”)Ch 2Promise-tracking + judge layer
Beat re-landed in different wordsCh 5Redundancy product (out of scope here)
Narrator over-explaining the theme (an AI tell)Ch 8Dev-Edit product (out of scope here)

Two honest notes. First, this run is the mechanical layer only — the entity-state contradiction detections (the character, world, object, co-presence and promise items above) are produced by the audit's Ledger-and-judge layer, which we didn't exercise in this reproducible-by-anyone run; we'll publish a second benchmark that does. Second, the deterministic tense pass is deliberately over-inclusive: it raised five candidates and left the decision to the review step, which is why it appears here as a limitation rather than a clean catch. We would rather tell you that than hide it.

This is why we can guarantee it. Because every finding is anchored to a line we can point to, we put money on it: if a top-severity finding in your report is ever shown to be factually wrong — the quote isn't in your book, or it misreads the passage — that report is free. The benchmark above is the mechanism that makes that promise safe to offer.

See what your manuscript's report would look like →