№ 01 · RCA Document Library
Digitise your paper archive into a table you can search and audit.
The Document Library is a self-hosted app for digitising and cataloguing documents. It OCRs your PDFs and scans, extracts the fields you define, and gives every value a confidence score you can actually interrogate. The result is a clean, searchable table of your entire archive.
How it works
Four stepsLoad your documents
PDFs and scans, digital or paper-sourced.
Define templates for the fields you want
The template builder tells you which fields depend on layout and which generalise.
Extraction runs with per-field confidence
Confidence is built from four visible factors: OCR quality, rule match, type check and shape check, so there is no black-box score.
Search, review, export
Review the flagged fields and export. Ground truth records are append-only, so your history never breaks.
Why teams pick it over digitisation vendors
All claims verifiable- 96.1% field accuracy on OCR-only extraction in the V2 benchmark (76 fields). That number is OCR only, with no AI assistance.
- Offline-first. OCR, extraction, search and export all run without internet. Your documents never leave your machines.
- Locked-down networking. The app enforces a runtime allowlist of exactly three hosts, and only if you turn AI assistance on.
- Optional AI scoring layer. If you add a provider key, AI validates extractions as a separate result. It never overwrites what the OCR engine found.
- Australian formats understood, including Medicare numbers.
- USD $0.10 per document extracted. No subscription and no upfront fee, a fraction of a standard digitisation contract.
What it costs
Consumption based| Stage | What you get | Price |
|---|---|---|
| Evaluation | A trial run on your own document types, on your hardware or on synthetic stand-ins. You see real extraction results before you commit to anything. | Free |
| Extraction | Charged per document extracted. Covers OCR, extraction, confidence scoring and cataloguing. | USD $0.10 per document |
| Licence and subscription | None. The app runs on your own hardware, single desktop or self-hosted Docker server, and you pay only for the documents you extract. | $0 |
For comparison: Australian scanning bureaus typically charge 8 to 25 cents per page for scanning alone, before any data extraction or cataloguing. Prices are in US dollars, with no subscription and no upfront fee.
Extraction accuracy depends on your templates and the condition of your documents. 96.1% is our published benchmark, not a promise about your corpus. That is why evaluation on your own document types comes first, before any commitment.
Run it on your documents.
Request an evaluation. It runs on your hardware or on synthetic data, so nothing sensitive changes hands.