№ 01 · RCA Document Library
From a box of paper to a table you can query.
The Document Library is a self-hosted app for digitising and cataloguing documents. It OCRs your PDFs and scans, extracts the fields you define, and gives every value a confidence score you can actually interrogate. The result is a clean, searchable table of your entire archive.
How it works
Four steps · no black boxLoad your documents
PDFs and scans, digital or paper-sourced.
Define templates for the fields you want
The template builder tells you which fields depend on layout and which generalise.
Extraction runs with per-field confidence
Confidence is built from four visible factors: OCR quality, rule match, type check and shape check. No black box scores.
Search, review, export
Review the flagged fields and export. Ground truth records are append-only, so your history never breaks.
Why teams pick it over digitisation vendors
The claims · all verifiable- 96.1% field accuracy on OCR-only extraction in the V2 benchmark (76 fields). No AI required for that number.
- Offline-first. OCR, extraction, search and export all run without internet. Your documents never leave your machines.
- Locked-down networking. The app enforces a runtime allowlist of exactly three hosts, and only if you turn AI assistance on.
- Optional AI scoring layer. If you add a provider key, AI validates extractions as a separate result. It never overwrites what the OCR engine found.
- Australian formats understood, including Medicare numbers.
- Costs a fraction of a standard digitisation contract, and you own the setup.
What it costs
Quote-based · anchors published| Stage | What you get | Price |
|---|---|---|
| Evaluation | A trial run on your own document types, on your hardware or on synthetic stand-ins. You see real extraction results before you commit to anything. | Free |
| Extraction | Per-document fee once you run at volume. Covers OCR, extraction, confidence scoring and cataloguing. | 10c per document |
| Licence and setup | Scoped to your deployment: single desktop or self-hosted Docker server, template building, and handover. | Quote |
For comparison: Australian scanning bureaus typically charge 8 to 25 cents per page for scanning alone, before any data extraction or cataloguing. Tell us your document volume and we will give you a full quote.
Extraction accuracy depends on your templates and the condition of your documents. 96.1% is our published benchmark, not a promise about your corpus. That is why evaluation on your own document types comes first, before any commitment.
Run it on your documents.
Request an evaluation. It runs on your hardware or on synthetic data, so nothing sensitive changes hands.