Evidence · run logs, not brochures
We publish our misses.
Most vendors quote one accuracy number. We publish run logs: every field on every document, including the ones the models got wrong. When a model fails a document, it is scored zero and recorded in the results.
Latest live trial · 05/08/2026
Run rca-library-trial-20-w3 · 20 documents20 documents, five fields each, ground truth known before the run. Sonnet 5 and Opus 5 completed all 20. Haiku 4.5 completed 18 and hit its output limit on two dense flow sheets, so those two score zero in the table. Document type identification is our current weak spot.
Total API spend for this run: A$4.42
- Haiku 4.5 · doc-0042: anthropic: output_truncated - reply hit the max_tokens output limit
- Haiku 4.5 · doc-0044: anthropic: output_truncated - reply hit the max_tokens output limit
| Field | Sonnet 5 | Haiku 4.5 | Opus 5 | Result |
|---|---|---|---|---|
| Entity | 20/20 | 18/20 | 20/20 | strong |
| Amount | 20/20 | 20/20 | 20/20 | clean |
| Document date | 19/20 | 18/20 | 20/20 | strong |
| Reference number | 19/20 | 16/20 | 16/20 | needs work |
| Document type | 14/20 | 15/20 | 14/20 | needs work |
| Overall field checks | 92/100 | 87/100 | 90/100 | honest |
V2 OCR benchmark · no AI assistance
13 documents · 76 fields96.1% field accuracy (73/76) on OCR-only extraction across 76 fields. No AI, no provider key, nothing leaves the machine. The per-template breakdown is on the right, misses included.
Dataset: tests/bench/extract-v2 (bundled synthetic self-test)
Generated 04/06/2026
| Template | Correct | Accuracy |
|---|---|---|
| finance-invoice | 15/16 | 93.8% |
| finance-receipt | 9/10 | 90.0% |
| clinical-lab-report | 14/14 | 100.0% |
| insurance-claim | 9/10 | 90.0% |
| insurance-eob | 7/7 | 100.0% |
| finance-bank-statement | 5/5 | 100.0% |
| identity-passport | 6/6 | 100.0% |
| identity-drivers-license | 5/5 | 100.0% |
| generic-keyvalue | 3/3 | 100.0% |
| All templates | 73/76 | 96.1% |
Methodology
How these numbers are made- Ground truth is known before the run. Every trial document comes from our generator, so the right answer for every field exists before any model sees the page.
- These tables render from the raw JSON logs. When we run a new trial, we drop the result file in and the page updates. Nothing is typed in by hand.
- Spend is published per run. The 05/08/2026 trial cost A$4.42 in API fees.
- Failures score zero. A model that does not answer does not get excused from the denominator.
Want this run against your document types?
Ask for a benchmark pack. Same methodology, your formats, ground truth included.