OCR for historical archives
A comparison of text recognition engines
Which OCR engine should you point at a box of archive scans? We built a fixed corpus of Los Angeles stage and picture-palace ephemera, generated reference transcripts, and measured every engine we could install against every document. This is that work, in order.









Nine items on one theme — the Los Angeles stage and its picture palaces, c. 1910–1945. Click any of them to see what the engines were actually given. Provenance and rights for every item are in the corpus reference.
The challenge
Why OCR matters for archival work, and what makes historical documents difficult to digitise.
The engines
What exists, how each one works, and the difference between pattern matching, neural recognition and layout understanding.
The corpus
Nine real documents, chosen to span the failure modes that matter in archive digitisation rather than to flatter anything.
The results
What the measurements say: which engine to reach for, and where each one falls down.
What follows
The findings that survived, and the questions still open.
Thousands of documents await digitisation, and every one of them is different
Faded ink and poor contrast
Aged paper and aged ink converge. Letters blend into the ground, and the threshold that recovers one part of a page destroys another.
Decorative typography
Historic mastheads and showbills use hand-lettered, arched and ornamented letterforms that no character model has been trained on.
Complex page layouts
Multi-column articles, mixed sizes, and text flowing around illustration. Finding the words and reading them in the right order are separate problems, and most engines solve only the first.
Volume
Transcribing a collection by hand is not a plan. Whatever runs has to run unattended, and has to say something useful when it fails.
What we measured, and what each one is
Free · local
tesseract
Traditional pattern-matching OCR, and the one everybody already has. Fast, predictable, and entirely dependent on being handed a well-prepared image.
Free · local
tesseract-auto
The same engine with contrast and page-segmentation chosen per image instead of left at the default. Tuning the engine you have turned out to beat switching to a different one.
Free · local
easyocr
A neural detector and recogniser. Finds text regions the others miss, particularly on textured and low-contrast grounds, then struggles to put them in order.
Free · local
paddle
Baidu’s detection-and-recognition pipeline. Strong word coverage on dense images, and the best local reader of small type on a busy sheet.
Free · local
marker
A document-to-Markdown converter rather than an OCR engine: it models the page as structure — headings, columns, tables — and reads within it.
Free · local · routing
auto-local
Not a standalone engine. It runs every local engine eligible for the file, scores the transcripts with no reference text, and keeps the best one. Per file, not per batch.
API · paid · ceiling reference
claude
A vision model, and the source of the reference transcripts every other engine is scored against — including its own. Its column is self-consistency, not accuracy, and it is excluded from every ‘best local’ comparison on this page.
Five documents, and what each one breaks
Document 1 of 5
Inside Facts of Stage and Screen
LA theatrical trade weekly, 31 May 1930 · JPEG
Dense multi-column newsprint under a heavy display masthead, on aged low-contrast paper. Every local engine finds most of the words and then reads them in the wrong order, so recall holds up while character similarity collapses. Finding the text and understanding the page are different problems, and this is the document that separates them.
Character similarity / word recall
tesseract- 0.18 / 0.85
tesseract-auto- 0.35 / 0.91
easyocr- 0.03 / 0.71
paddle- 0.21 / 0.94
marker- 0.39 / 0.90
auto-local- 0.39 / 0.90
claude*- 1.00 / 1.00
Bold is best local. * ceiling reference, not a measurement.
Document 2 of 5
The great Victorina Troupe
Vaudeville chromolithograph, c. 1914 · JPEG
Curved and arched hand-lettered display type, colour on colour, with caption text at the bottom in a size the scan barely holds. Nothing local reads it — plain Tesseract returns essentially nothing at all. The widest gap in the corpus between what a local engine manages and what a vision model does.
Character similarity / word recall
tesseract- 0.00 / 0.00
tesseract-auto- 0.02 / 0.01
easyocr- 0.05 / 0.39
paddle- 0.13 / 0.39
marker- 0.01 / 0.00
auto-local- 0.13 / 0.39
claude*- 0.73 / 0.98
Bold is best local. * ceiling reference, not a measurement.
Document 3 of 5
Carthay Circle Theatre, reverse
Tichnor linen postcard, c. 1930–45 · PNG
Clean printed type, awkwardly arranged: rotated ninety degrees, around a faded rubber stamp and a line of handwriting, with large empty areas between. Auto-configured Tesseract comes within a few points of the ceiling here — free, offline and fast. This is the case where reaching for an API is hard to justify.
Character similarity / word recall
tesseract- 0.59 / 0.50
tesseract-auto- 0.91 / 0.71
easyocr- 0.79 / 0.62
paddle- 0.28 / 0.62
marker- 0.66 / 0.62
auto-local- 0.91 / 0.71
claude*- 0.96 / 0.96
Bold is best local. * ceiling reference, not a measurement.
Document 4 of 5
Grauman’s Chinese Theatre
Tichnor linen postcard, c. 1930–45 · JPEG
A colour halftone where the marquee lettering sits right at the edge of legibility and competes with the image around it. The local engines cluster on character similarity but recover only about half the words, and Marker’s layout model reads the whole card as a figure and returns almost nothing.
Character similarity / word recall
tesseract- 0.74 / 0.54
tesseract-auto- 0.86 / 0.54
easyocr- 0.84 / 0.54
paddle- 0.78 / 0.31
marker- 0.09 / 0.00
auto-local- 0.78 / 0.31
claude*- 1.00 / 1.00
Bold is best local. * ceiling reference, not a measurement.
Document 5 of 5
The King of Kings souvenir programme
Roadshow programme, 1927 · 21-page PDF
The only multi-page item, and the reason PDF support is not a detail: EasyOCR and PaddleOCR cannot open it at all. Tesseract and Marker both manage strong recall on the typeset body text, across tinted grounds, script headings and drop caps.
Character similarity / word recall
tesseract- 0.45 / 0.95
tesseract-auto- 0.45 / 0.95
easyocr- N/A
paddle- N/A
marker- 0.66 / 0.93
auto-local- 0.66 / 0.93
claude*- 0.89 / 0.97
Bold is best local. * ceiling reference, not a measurement.
Choosing per file beats choosing an engine
Averaged over the whole corpus. Character similarity is order-sensitive and word recall is not, so they are always reported as a pair — the gap between them is where the interesting behaviour lives.
Every engine against every document
| Fixture | tesseract | tesseract-auto | claude* | easyocr | paddle | marker | auto-local |
|---|---|---|---|---|---|---|---|
carthay-circle-postcard-back.png | 0.59/0.50 | 0.91/0.71 | 0.96/0.96 | 0.79/0.62 | 0.28/0.62 | 0.66/0.62 | 0.91/0.71 |
carthay-circle-premiere.jpg | 0.00/0.00 | 0.49/0.86 | 0.89/0.93 | 0.87/0.71 | 0.09/0.00 | 0.76/0.86 | 0.49/0.86 |
graumans-chinese-theatre.jpg | 0.74/0.54 | 0.86/0.54 | 1.00/1.00 | 0.84/0.54 | 0.78/0.31 | 0.09/0.00 | 0.78/0.31 |
hollywood-music-box-playbill-1926.tif | 0.19/0.46 | 0.19/0.46 | 0.90/0.97 | 0.09/0.39 | 0.39/0.90 | 0.28/0.74 | 0.28/0.74 |
inside-facts-1930-cover.jpg | 0.18/0.85 | 0.35/0.91 | 1.00/1.00 | 0.03/0.71 | 0.21/0.94 | 0.39/0.90 | 0.39/0.90 |
inside-facts-1930-page-six.jpg | 0.45/0.90 | 0.37/0.89 | 0.94/1.00 | 0.04/0.78 | 0.10/0.94 | 0.40/0.94 | 0.40/0.94 |
kar-mi-troupe-poster.jpg | 0.00/0.00 | 0.02/0.01 | 0.73/0.98 | 0.05/0.39 | 0.13/0.39 | 0.01/0.00 | 0.13/0.39 |
kinema-theater-ad-1920.tif | 0.23/0.40 | 0.04/0.38 | 0.76/0.99 | 0.10/0.66 | 0.22/0.93 | 0.75/0.88 | 0.75/0.88 |
king-of-kings-souvenir-1927.pdf | 0.45/0.95 | 0.45/0.95 | 0.89/0.97 | N/A | N/A | 0.66/0.93 | 0.66/0.93 |
| Average | 0.31/0.51 | 0.41/0.63 | 0.90/0.98 | 0.35/0.60 | 0.28/0.63 | 0.44/0.65 | 0.53/0.74 |
Generated from evaluation/benchmark.csv at build time.
Bold marks the best local backend per fixture.
claude generated the ground truth and is a ceiling reference rather than a score.
The same numbers, drawn

Claude sets a ceiling, and is not a competitor
It generated the reference transcripts every other engine is scored against, including its own. Reading its column as a win is reading the wrong thing. What it does tell you is roughly how much of each document is legible at all.
No single local engine is strong across the corpus
Different engines win on different documents, and the winners are not close to each other. There is no setting you can choose once and be right about, which is the finding that decided the architecture.
Choosing per file closes most of the gap
auto-local runs the engines eligible for each file, scores the transcripts without a reference, and keeps the best. It beats every individual local engine on both metrics, entirely offline and free — and it is still losing ground to an oracle that knows the answer.
Which engine for which material?
- Clean printed captions, postcards, labels
tesseract-autotheneasyocr- Close to the ceiling, free, offline and fast
- Multi-column newsprint and magazines
claudethenmarker- Reading order is the whole problem, and only layout understanding solves it
- Decorative, hand-lettered, ornamental type
claude- Nothing local reads it
- Multi-page PDFs
markerthentesseract- EasyOCR and PaddleOCR cannot read PDFs at all
- Mixed material, run unattended
auto-local- Runs the viable local engines and keeps the best transcript
- Anything confidential
auto-localthentesseract-auto- Everything stays on the machine; no API call
What is still open
The selection heuristic leaves accuracy on the table
auto-local’s scorer loses ground against an oracle that knows which transcript is best. That is a scoring weight rather than an engine limitation, and it is the cheapest accuracy work available.
The auto-configuration bands are overfitted
They were fitted against most of this corpus and need re-fitting on a wider set before they can be trusted elsewhere.
The ground truth is machine-generated
Human-checked reference transcripts would remove the benchmark’s central weakness, and would let Claude be measured rather than assumed.
Nothing here generalises beyond archival print of this era
Same period, same condition, same script. Whether the pipeline holds up on other scripts is untested; an open Armenian Tesseract model exists and would make the cheapest first experiment.
Rights, before anything else
Would the rights attached to a given collection permit sending its scans to a third-party service at all? Where the answer is no, the local path is not a preference, it is the only option.