The research in depth
Tetrak OCR exists because of a benchmark, not the other way round. This is that work end to end: why we ran it, what we measured, what it found, the finding we got wrong, and the four mechanisms the evidence forced into the product.
Read it in order and it is an argument. Jump around with the contents on the right and it is a reference.
This is the detailed version. The walkthrough covers the same ground in five beats, with the corpus and the results drawn out, and is the better place to start if you have not read either.
Why run a benchmark at all
A real archive is not a stack of uniform pages. The nine items in this corpus include a linen postcard, a chromolithograph poster, two pages of 1930 newsprint, a bilevel TIFF of a theatre playbill and a twenty-one-page souvenir programme. They differ in typography, contrast, layout and physical condition, and they were digitised by different people at different times.
The obvious way to build an OCR tool for that material is to pick a good engine, configure it well, and run it over the batch. Every vendor comparison, every “best OCR library” article, and most of our own instinct pointed that way.
We could not find evidence that it was the right shape. So before designing anything we built a fixed corpus, generated reference transcripts, and measured every engine we could install against every document. The result decided the architecture — and, in one case, unbuilt a feature that had already shipped.
What we measured
Nine items of Los Angeles stage and picture-palace ephemera, c. 1910–1945, chosen to span the failure modes that matter in archive digitisation rather than to flatter anything. Provenance and rights for every item are in the corpus reference.
Each transcript is scored on two numbers, both against normalised text — lowercased, whitespace collapsed.
Character similarity compares the transcript to the reference as a sequence. It is order-sensitive: text read in the wrong order scores badly even when every word is present.
Word recall is the share of expected words appearing anywhere in the transcript. It is order-insensitive, and it ignores extra words — an engine that invents text is not penalised.
They are always reported as a pair, because they fail differently and the gap between them is where the interesting behaviour lives:
| Pattern | Means |
|---|---|
| High recall, low similarity | The words were found and read in the wrong order — the signature of a multi-column page read as one column |
| High similarity, low recall | What was read was read correctly, but some text was never seen — faint lettering, ornament, low contrast |
| Both low | The engine could not read the page |
On our deliberately unruly corpus, we found that no single tool performed well on all items, and gained valuable insight into where each tool stood out.
The weakness at the centre of the method
Reference transcripts were generated by Claude (claude-opus-4-8) and are
committed under evaluation/ocr/corpus/expected/.
The reference is committed rather than regenerated per run precisely because of that instability; a moving reference would make run-to-run comparison meaningless. Human-checked ground truth would fix this properly, and it is the single highest-value improvement available to this work.
What the numbers support: relative comparison between local engines on this kind of material, and relative speed. The metrics are deterministic, the corpus is fixed and public, the reference is committed, so the local columns reproduce exactly.
What they do not support: any claim about Claude’s absolute accuracy, or generalisation beyond archival print of this era and condition.
The results
Character similarity / word recall, nine fixtures, run on an Apple MPS machine
so auto-local had Marker available. Bold marks the best local engine per
fixture; N/A means the engine cannot read PDFs.
| Fixture | tesseract | tesseract-auto | claude* | easyocr | easyocr-hy | paddle | paddle-vl | marker | vision | auto-local |
|---|---|---|---|---|---|---|---|---|---|---|
bancroft-magician-poster.jpg | 0.07/0.00 | 0.28/0.00 | 0.72/0.77 | 0.40/0.14 | 0.17/0.00 | 0.55/0.32 | 0.31/0.14 | 0.17/0.00 | 0.57/0.41 | 0.57/0.41 |
carthay-circle-postcard-back.png | 0.59/0.50 | 0.93/0.79 | 0.98/0.92 | 0.79/0.62 | 0.49/0.08 | 0.28/0.62 | 0.68/0.79 | 0.66/0.62 | 0.65/0.83 | 0.93/0.79 |
carthay-circle-premiere.jpg | 0.00/0.00 | 0.51/0.86 | 0.89/0.93 | 0.87/0.71 | 0.74/0.29 | 0.09/0.00 | 0.88/0.86 | 0.76/0.86 | 0.89/0.93 | 0.51/0.86 |
chinese-theatre-triptych-postcard.jpg | 0.44/0.40 | 0.44/0.40 | 1.00/1.00 | 0.50/0.67 | 0.46/0.00 | 0.79/0.73 | 0.83/0.87 | 0.45/0.47 | 0.66/0.93 | 0.45/0.47 |
graumans-chinese-theatre.jpg | 0.66/0.44 | 0.79/0.44 | 0.88/0.56 | 0.72/0.50 | 0.61/0.06 | 0.74/0.19 | 0.72/0.31 | 0.10/0.00 | 0.73/0.38 | 0.73/0.38 |
greek-theatre-night.jpg | 0.90/0.75 | 0.90/0.75 | 1.00/1.00 | 0.95/0.75 | 0.79/0.00 | 0.91/0.50 | 0.99/0.92 | 0.74/1.00 | 0.98/1.00 | 0.74/1.00 |
hollywood-boulevard-east.jpg | 0.63/0.33 | 0.62/0.33 | 0.93/0.83 | 0.72/0.44 | 0.65/0.17 | 0.64/0.44 | 0.65/0.39 | 0.11/0.00 | 0.68/0.39 | 0.68/0.39 |
hollywood-boulevard-west.jpg | 0.42/0.27 | 0.42/0.27 | 0.96/0.83 | 0.57/0.30 | 0.42/0.17 | 0.58/0.43 | 0.49/0.27 | 0.08/0.00 | 0.53/0.43 | 0.42/0.27 |
hollywood-music-box-playbill-1926.tif | 0.52/0.46 | 0.57/0.53 | 0.99/0.96 | 0.50/0.39 | 0.34/0.26 | 0.78/0.91 | 0.88/0.90 | 0.47/0.74 | 0.69/0.81 | 0.47/0.74 |
inside-facts-1930-cover.jpg | 0.68/0.85 | 0.90/0.91 | 1.00/0.99 | 0.41/0.71 | 0.65/0.50 | 0.50/0.94 | 0.98/0.93 | 0.75/0.90 | 0.63/0.88 | 0.75/0.90 |
inside-facts-1930-page-six.jpg | 0.94/0.91 | 0.94/0.91 | 0.99/0.95 | 0.30/0.78 | 0.61/0.60 | 0.30/0.97 | 0.96/0.95 | 0.74/0.93 | 0.59/0.89 | 0.74/0.93 |
kar-mi-troupe-poster.jpg | 0.01/0.00 | 0.09/0.01 | 0.99/0.98 | 0.12/0.39 | 0.10/0.00 | 0.27/0.38 | 0.29/0.44 | 0.03/0.00 | 0.29/0.46 | 0.29/0.46 |
kinema-theater-ad-1920.tif | 0.45/0.40 | 0.45/0.40 | 0.75/0.98 | 0.42/0.67 | 0.43/0.33 | 0.49/0.93 | 0.74/0.92 | 0.93/0.87 | 0.88/0.91 | 0.93/0.87 |
king-of-kings-souvenir-1927.pdf | 0.91/0.94 | 0.91/0.94 | 0.93/0.97 | N/A | N/A | N/A | 0.89/0.90 | 0.87/0.92 | N/A | 0.87/0.92 |
over-the-fence-poster.jpg | 0.19/0.14 | 0.40/0.18 | 0.77/0.82 | 0.49/0.23 | 0.27/0.14 | 0.64/0.45 | 0.39/0.36 | 0.19/0.00 | 0.54/0.45 | 0.64/0.45 |
parlor-match-poster.jpg | 0.00/0.00 | 0.14/0.03 | 1.00/1.00 | 0.67/0.44 | 0.24/0.00 | 0.68/0.31 | 0.88/0.72 | 0.76/0.75 | 0.79/0.66 | 0.79/0.66 |
thurston-magician-poster.jpg | 0.08/0.04 | 0.26/0.26 | 0.99/0.96 | 0.74/0.52 | 0.12/0.00 | 0.74/0.44 | 0.73/0.63 | 0.72/0.52 | 0.62/0.52 | 0.72/0.52 |
| Average | 0.44/0.38 | 0.56/0.47 | 0.93/0.91 | 0.57/0.52 | 0.44/0.16 | 0.56/0.54 | 0.72/0.66 | 0.50/0.50 | 0.67/0.68 | 0.66/0.65 |
Generated from evaluation/benchmark.csv at build time.
Bold marks the best local backend per fixture.
claude generated the ground truth and is a ceiling reference rather than a score.
| Fixture | tesseract | tesseract-auto | claude* | easyocr | easyocr-hy | paddle | paddle-vl | marker | vision | auto-local |
|---|---|---|---|---|---|---|---|---|---|---|
bancroft-magician-poster.jpg | 0.5 | 0.3 | 3.9 | 3.0 | 4.0 | 10.9 | 23.1 | 8.9 | 0.9 | 15.1 |
carthay-circle-postcard-back.png | 0.4 | 0.5 | 6.2 | 0.4 | 0.4 | 7.2 | 21.2 | 5.7 | 0.6 | 16.8 |
carthay-circle-premiere.jpg | 0.2 | 0.5 | 2.4 | 0.3 | 0.3 | 6.8 | 18.1 | 5.6 | 0.5 | 15.3 |
chinese-theatre-triptych-postcard.jpg | 0.3 | 0.3 | 2.2 | 0.3 | 0.3 | 6.0 | 13.7 | 4.4 | 0.5 | 13.2 |
graumans-chinese-theatre.jpg | 0.3 | 0.3 | 2.4 | 0.3 | 0.3 | 6.3 | 16.6 | 5.4 | 0.5 | 13.5 |
greek-theatre-night.jpg | 0.3 | 0.3 | 2.4 | 0.3 | 0.3 | 4.3 | 15.1 | 5.4 | 0.5 | 12.3 |
hollywood-boulevard-east.jpg | 0.3 | 0.3 | 2.8 | 0.3 | 0.3 | 7.4 | 11.0 | 5.6 | 0.5 | 15.5 |
hollywood-boulevard-west.jpg | 0.3 | 0.3 | 3.5 | 0.3 | 0.3 | 6.4 | 17.9 | 6.8 | 0.5 | 14.8 |
hollywood-music-box-playbill-1926.tif | 1.0 | 1.1 | 28.1 | 5.7 | 7.8 | 42.6 | 170.3 | 274.7 | 1.2 | 325.3 |
inside-facts-1930-cover.jpg | 2.2 | 2.2 | 27.2 | 7.8 | 9.7 | 57.5 | 221.1 | 286.8 | 2.2 | 349.6 |
inside-facts-1930-page-six.jpg | 4.9 | 4.9 | 64.0 | 17.7 | 21.5 | 132.4 | 612.5 | 308.0 | 4.2 | 461.4 |
kar-mi-troupe-poster.jpg | 0.3 | 0.4 | 11.2 | 0.7 | 1.0 | 9.2 | 27.0 | 2.9 | 0.8 | 16.8 |
kinema-theater-ad-1920.tif | 0.8 | 0.8 | 17.3 | 2.7 | 3.8 | 26.3 | 126.5 | 23.2 | 1.1 | 58.3 |
king-of-kings-souvenir-1927.pdf | 35.5 | 35.2 | 150.7 | — | — | — | 1128.7 | 423.3 | — | 431.6 |
over-the-fence-poster.jpg | 0.3 | 0.3 | 3.0 | 0.3 | 0.3 | 6.5 | 14.7 | 3.2 | 0.6 | 12.9 |
parlor-match-poster.jpg | 0.4 | 0.3 | 3.5 | 0.4 | 0.6 | 7.9 | 25.7 | 7.0 | 0.8 | 19.3 |
thurston-magician-poster.jpg | 0.3 | 0.3 | 3.2 | 0.3 | 0.4 | 6.3 | 18.1 | 6.4 | 0.6 | 19.4 |
| Total seconds | 48.2 | 48.3 | 333.9 | 40.8 | 51.4 | 344.0 | 2481.4 | 1383.3 | 16.1 | 1811.0 |
Generated from evaluation/benchmark.csv at build time.
claude generated the ground truth and is a ceiling reference rather than a score.

Both the table and the chart are generated from evaluation/ocr/benchmark.csv at
build time, so neither can drift from the data.
What the numbers say
No engine wins across the board
This is the headline, and it is why the average column is the least useful part of the table. Five different engines win across the nine fixtures:
- Marker takes the newsprint cover (0.39), the Kinema ad (0.75) and the PDF (0.66)
- Retuned Tesseract takes the postcard reverse (0.91) and Grauman’s (0.86)
- PaddleOCR takes the playbill (0.39) and the poster (0.13)
- EasyOCR takes the Carthay Circle front (0.87)
- Plain Tesseract takes the interior newsprint page (0.45)
An earlier version of this page added “and none wins more than twice”. That is no longer true — Marker takes three, both TIFFs having landed in territory it handles well. The claim that matters is unchanged: no engine wins most of them, and the winner is not predictable from the average.
Route by document, not by average.
The metrics disagree, and that is the point
On newsprint every local engine pairs high word recall (0.71–0.94) with near-zero character similarity (0.03–0.45). The words are all there; the columns are interleaved. Character similarity is order-sensitive and word recall is not, so the gap between them is the layout failure, quantified.
PaddleOCR recovers 94% of the words on the Inside Facts front page and scores 0.21 on similarity. For a human reader that output is unusable; for full-text search it is adequate. Which number matters depends on whether you are building a reading edition or a search index, and that is your question rather than ours — which is why we never collapse the two into one score.
Tuning beat switching
tesseract-auto averages 0.41/0.63, ahead of EasyOCR at 0.35/0.60 — a
deep-learning engine beaten by Tesseract with better per-image settings.
It has not always been. It previously scored 0.22 and lost to plain Tesseract.
The fix was re-fitting the page-segmentation bands in analyse_image(), which
had been inherited from a parent project and were sending dense pages to the
sparse-text mode. Carthay Circle’s front went 0.00/0.00 → 0.49/0.86 on that
change alone.
The same engine, the same image, a different page-segmentation mode. Reaching for a heavier engine is the expensive reflex — gigabytes of weights, minutes of runtime, sometimes an API bill — and on this evidence it is often the wrong first move.
…but the first tuning did not generalise
Two TIFFs were added after analyse_image()’s thresholds were first fitted by
hand, which made them the only held-out data here. Auto-configuration was worth
nothing on either — it gained nothing on the playbill and lost 0.19 on the
Kinema ad, against a gain of +0.15 on the seven it had been fitted to. On the
fitted set it gained on five fixtures, once by +0.49. That is what overfitting
looks like, and with seven images against four contrast bands and three PSM
bands there was ample freedom for it to happen.
The Kinema failure was diagnosable rather than mysterious. Its frame is 42.9% dark because a large illustration fills half of it, which put it in the “image-dominated, no page structure worth analysing” band and selected PSM 6, a single uniform block. But it does have structure: a narrow programme column in tiny type beside the headline block. PSM 3 read it at 0.23; PSM 6 collapsed it to 0.04.
analyse_image()’s own docstring anticipated the risk — “it cannot tell ink
from imagery” — and reasoned that a dark frame has little structure to analyse
anyway. That held for the night-scene postcards the band was fitted against. It
does not hold for an illustrated page that still has columns, and a histogram
cannot tell those two apart.
Re-fitting them, and what it cost
On 27 September 2026 the bands stopped being hand-written. They are now chosen
by evaluation/ocr/calibration/fit.py,
which scores all 42 (contrast, psm) pairs against every raster fixture and
picks the band edges that minimise mean regret — how far each fixture falls
below the best it could have scored. Regret rather than mean score, because the
corpus spans a 350× range in transcript length and a mean hands the objective
to the longest documents. The current values are on
the routing reference, read from the module rather
than typed into the page.
Five of the nine tesseract-auto figures moved, and all five moved up. The
Kinema ad now matches plain Tesseract exactly instead of collapsing; the
Inside Facts inner page, which had been quietly losing 0.09 to the contrast
band, gains it back; the Carthay Circle premiere, the playbill and the Kar-Mi
poster each gain a little. The postcard reverse, Grauman’s and the Inside
Facts cover were already on their fitted cell and did not move, and the PDF
cannot — the bands never run on that path. Nothing scores below plain anywhere
in the corpus now, where two fixtures did.
Repairing the Kinema ad meant spending the held-out data. The PSM-6 edge was unconstrained anywhere between 25.3% and 56.1% dark on the six fitted fixtures, and the Kinema ad sits at 42.9% — inside that gap. No fit that could not see it would ever have moved the edge off it. So it was promoted into the fit set, and for a day there was nothing held out.
That left leave-one-out as the only generalisation measure, and it was not reassuring: refitting without each fixture in turn and scoring it with the bands that result gave a mean regret of 0.20 against 0.04 in sample. The conclusion drawn at the time was that eight fixtures was the binding constraint, and that more scans would fix it.
Doubling the corpus, and what it actually showed
On 28 September the corpus went from nine fixtures to seventeen. Eight rasters — four linen postcards and four lithographed theatre posters — had been sourced in August specifically to widen calibration and had sat unusable ever since, because none of them had a reference transcript. Generating those, and then checking every transcript in the corpus by hand for the first time, restored a genuinely held-out set of four.
The check was worth more than the widening. Five of the seventeen references were wrong. Two omitted a printer’s imprint entirely; one read a fragment of musical notation as the number 18; two were outright fabrications on the Kar-Mi poster, where “drinking water through a tube” had become “drinking water successfully”. And one was worth 0.25 character similarity on its own: the Inside Facts inner page had been reflowed into continuous paragraphs, joining 27 end-of-line hyphenations that the newspaper actually prints. Every engine had been penalised for reading the page correctly.
A second Kinema turned up immediately, in material nothing had been fitted
to. The Greek Theatre postcard is 69.4% dark — a night scene — so the bands
read it as image-dominated and sent it to PSM 6, scoring 0.37. Plain Tesseract
reads it at 0.90. The dark_pct heuristic fails at the top of its range as
well as in the gap the Kinema ad sat in, which the original corpus could not
have shown because the Kinema ad was the only dark-framed fixture in it.
And generalisation got worse, not better. Leave-one-out regret is 0.265 over twelve fitted fixtures, against 0.20 over eight. More data made the problem harder, because the material is now genuinely varied: postcards, newsprint, a playbill and chromolithographs do not share one mapping from two pixel statistics to a configuration. Nor does spending more bands help — the figure is identical at four contrast bands and at five. Neither the corpus size nor the model’s complexity is the limit. Two features are.
That is a more useful answer than the one the smaller corpus gave, and it is the opposite of it.
Note what did not happen when the two TIFFs were originally added: no other
engine’s ranking changed, and auto-local still routed the Kinema ad correctly by
picking Marker at 0.75. The quality score caught what the heuristic got wrong —
which is the argument for measuring transcripts rather than predicting from
pixels.
Fan-out wins, by less than it should
auto-local averaged 0.53/0.74 before 28 August 2026, the best local result on
both metrics, and its margin over tesseract-auto widened from 0.05 to 0.12 as
the corpus grew, because fan-out measures each transcript instead of predicting
from pixels. It now averages 0.57/0.76 — Vision joined the candidate pool
that day too, alongside the fix below, so the margin over tesseract-auto is
now 0.16.
The ceiling is higher, though closer than it was. An oracle taking the best
local engine on every fixture averages 0.61; auto-local reaches 0.57,
about 94% of it. Four fixtures accounted for the gap before the fix below; two
of them were a scoring problem, not an engine one, and are fixed:
| Fixture | Best local | Chose before | Chose after | Cost before → after |
|---|---|---|---|---|
carthay-circle-premiere.jpg | vision 0.89 | tesseract-auto 0.49 | marker 0.76 | 0.40 → 0.14 |
graumans-chinese-theatre.jpg | tesseract-auto 0.86 | paddle 0.78 | vision 0.83 | 0.08 → 0.03 |
hollywood-music-box-playbill-1926.tif | paddle 0.39 | marker 0.28 | marker 0.28 | 0.11 (held out; not fit to) |
inside-facts-1930-page-six.jpg | tesseract 0.45 | marker 0.40 | marker 0.40 | 0.05 (plain tesseract is not a candidate) |
The last row does not move, and cannot: auto-local never runs plain
tesseract, only tesseract-auto, so no scoring change closes a gap against
a backend the router does not have. The third is held out from fitting on the
same principle as the Tesseract auto-configuration bands above — “…but the
tuning does not generalise”.
The Carthay Circle case repays reading closely, because it shows the raw
quality score getting a candidate genuinely wrong, not just outranked by
length. Marker’s combined_score (0.1517) is fractionally the highest of
the four candidates — ahead of Vision’s (0.1405) and EasyOCR’s (0.1416) — even
though Marker’s character similarity (0.76) trails both. Under the old
ranking, quality × (0.5 + 0.5 × words/max_words), none of that mattered:
tesseract-auto emitted 77 words (mostly noise from the postcard’s textured
background) against 11 each for the other three, and word count alone
overrode all of them, promoting a transcript scoring 0.49 over one scoring
0.89 — the lowest-scoring candidate winning outright.
Raising the weight to 0.8 removes the length effect for this fixture
entirely — tesseract-auto’s own combined_score is the lowest of the four,
so no plausible weight would still pick it — which is enough to hand the
result to Marker instead. It cannot fix Marker’s ranking error, because that
error was never about length: evaluation/ocr/calibration/router_sweep.py
fits the weight by minimising mean regret against ground truth, the same
principle the Tesseract bands are held to. Under the corrected character
similarity (October 2026), mean regret over the fit set falls from 0.134 to
0.090 and 0.8 still wins, though one fixture, graumans-chinese-theatre, now
scores fractionally worse; under the metric in use when the weight was chosen,
the fall was 0.083 to 0.032 with no fixture worse.
the error taxonomy (product analytics) and
brief 010 (the router length weight) have the full account, including
why the residual 0.14 is now a different, smaller problem, inside
combined_score rather than the router.
A feature we measured and then removed
auto-local-fast ran only the top-ranked eligible engine instead of all of
them. On this machine that was always Marker, so it produced identical scores to
marker (0.42/0.61) at almost identical cost — strictly worse than an engine
already available:
| character similarity | seconds | |
|---|---|---|
tesseract-auto | 0.49 | 45 |
auto-local-fast | 0.42 | 1169 |
Lower accuracy and twenty-six times the runtime. Nor was that only a GPU-machine
problem: worked through every installed configuration, it never won. tesseract-auto
already occupied the niche fast mode was invented for and filled it better, so
the backend was removed, one day after it shipped.
Registering it as a backend rather than hiding it behind a flag is what turned a preference into a measurement, and then into a deletion.
Marker is uneven and slow
It wins the newsprint cover, the Kinema ad and the PDF, then collapses on Grauman’s (0.09/0.00), where its layout model reads the whole postcard as a figure and returns nearly nothing.
It also dominates the runtime. Of the roughly 65 minutes a full --all run
takes, about 50 are Marker. Dropping it from the fan-out pool would cost more
than it saves: it wins three of nine, including the Kinema ad at 0.75 where the
next best local result is plain Tesseract at 0.23.
The poster defeats everything local — and destabilises the reference
Best local character similarity is 0.13; retuned Tesseract manages 0.02.
Claude is the only thing that reads it, but how well is not a stable number. Across four runs against the same committed ground truth — which Claude itself generated — it scored 0.73, 0.24, 0.26 and 0.73. Word recall over the same four runs was 0.98, 0.96, 0.97 and 0.98.
That split is the finding. Claude recovers the same words every time and arranges them differently every time, because a chromolithograph has no single correct reading order for a transcript to be scored against. It is the same signal as the interleaved newsprint columns, arriving for a different reason.
Hand-lettered display type is where matching letter shapes stops working and understanding what the page is starts mattering. If your material looks like this, the local tools will not do — and this benchmark cannot tell you precisely how much better the alternative is.
A finding we retracted
An earlier version of this analysis claimed auto-local mis-ranked the PDF,
choosing Marker when Marker had scored 0.02 character similarity there.
That was wrong and has been retracted. The 0.02 was an artefact of truncated
ground truth — the Claude backend was calling the API with max_tokens=4096 and
silently returning partial transcripts for the twenty-one-page programme, so
Marker’s correct output was being compared against a fragment. Against the
corrected reference Marker scores 0.66 and is the best engine on that fixture.
The quality heuristic had been right all along.
It is recorded rather than quietly deleted because the failure mode is instructive: a truncation bug in the reference looked exactly like a routing bug in the system under test. The Claude backend now raises on truncation rather than returning partial text.
What the evidence built
Four findings, four mechanisms. Together they are auto-local, and together
they are the six stages between dropping a file in and getting a transcript out.

Fan out, and select per file
No engine wins, so the obvious design — pick one, configure it, apply it to the batch — is unavailable. Any single choice is wrong for most of the corpus.
auto-local runs every viable engine over each file and keeps the best
transcript. Selection happens per file, not per batch, which is the only
granularity at which the finding can be acted on. The cost is honest: several
engines per file, so it is the slowest local mode.
The engines run sequentially, deliberately. Each is heavy on CPU and memory and several hold model singletons; starting two or three at once trades predictable runtime for contention, and it hurts most on exactly the large batches where the saving would matter.
Auto-configure before reaching for a bigger engine
Tuning beat switching, so analyse_image() inspects each image before OCR and
picks contrast and page-segmentation mode from two cheap statistics:
| Signal | What it tells us | Effect |
|---|---|---|
| Standard deviation of pixel intensity | How tonally flat the scan is — faded ink on aged paper reads low | Contrast boost |
| Share of frame below luminance 128 | How much of the frame is not paper | Page-segmentation mode |
The second misleads by its name. dark_pct looks like a text-density measure
and is not: a night-scene postcard reads 56% dark because of the photograph.
That turned out to be useful anyway — a frame dominated by imagery has little
page structure for layout analysis to work with, however the darkness got there.
The heuristic works for a slightly different reason than the one it was reached
for, which is also why the hand-fitted bands failed on the Kinema ad above: the
signal is sound, and the edges drawn across it were not.
The exact bands are in the routing reference, read from the module at build time rather than written out here — they are fitted from a sweep and they move.
Score without a reference
The metrics disagree, and both need ground truth. At run time there is none, so whatever picks the winner has to judge a transcript it has never seen the answer to.
Two signals that fail differently, multiplied, then weighted by relative word count:
- Dictionary coverage — the share of tokens that are real words. Catches
character-level garbage like
e1rr0r. - Perplexity, via GPT-2 — how natural the text reads. Catches the real-word substitutions a spell check waves through.
The word-count weighting came straight out of an observed failure: on sparse images a conservative engine emitting eight clean words was beating one that recovered twenty-three noisier ones, because dictionary coverage is trivially 1.00 for a short clean phrase. Quality alone rewarded saying less. As the Carthay Circle case shows, that correction now overshoots in the other direction on at least one fixture.
Fail loudly
Some pages defeat everything local. For a tool running unattended over an
archive that is the dangerous case — not because it fails, but because a failed
transcript looks exactly like a successful one from outside. A .md file
appears in workspace/processed/, and nothing signals that it is noise.
So there is a quality floor of 0.10. Below it the scan is diverted to
workspace/triage/ with a manifest recording what every engine produced and
how each scored. It never reaches workspace/processed/.
The point is not that the pipeline handles hard documents. It is that you can
tell which ones it did not, which makes “what is still in triage/” a
meaningful question and an actionable queue.
What we chose not to build
No automatic escalation. The triage queue is passive. Forwarding failures to a vision model automatically would be easy, and would turn a local, free, offline tool into one making paid API calls on material the user may not have wanted to leave the machine. That decision stays with a person.
No single quality number. Character similarity and word recall are reported as a pair everywhere, because collapsing them hides the finding above.
No claim about Claude’s accuracy. It generated the ground truth it is scored against. It sets a ceiling reference and nothing more.
Choosing a tool for your own material
The practical version of “no engine wins”:
| Your material | Use | Second choice | Why |
|---|---|---|---|
| Clean printed captions, postcards, labels | tesseract-auto | easyocr | 0.91 against Claude’s 0.98. Free, offline, fast |
| Multi-column newsprint, magazines | claude | marker | Reading order is the whole problem, and only layout understanding solves it |
| Decorative, hand-lettered, ornamental type | claude | — | Nothing local reads it. Best local score is 0.13 |
| Multi-page PDFs | marker | tesseract | EasyOCR and PaddleOCR cannot read PDFs at all |
| Mixed material, unattended | auto-local | — | Runs the viable local engines, keeps the best transcript |
| Anything confidential | auto-local | tesseract-auto | Everything stays on the machine; no API call |
On the Carthay Circle postcard reverse, five points of character similarity
separate a free offline tool (tesseract-auto, 0.91) from a paid API call
(Claude, 0.96). For clean printed text that is the norm rather than the
exception, and it is the case where reaching for a vision model is hard to
justify.
If you do route decorative and multi-column material to Claude, a page runs roughly 1,500 input and 750 output tokens on this corpus:
| Model | Input | Output | Approx. per page |
|---|---|---|---|
| Haiku 4.5 | $1 / MTok | $5 / MTok | ~$0.005 |
| Opus 4.8 | $5 / MTok | $25 / MTok | ~$0.026 |
| Opus 4.8, Batch API | 50% off | 50% off | ~$0.013 |
A practical pattern: run auto-local over everything and send only what lands
in workspace/triage/ to Claude. On this corpus that would be the poster and
little else.
What is still open
- The selection heuristic loses 0.07 average against an oracle that knows the answer. That is a scoring weight, not an engine limitation, and it is the cheapest accuracy work available.
- The auto-configuration bands are overfitted to seven of nine fixtures and need re-fitting against a wider set.
- Ground truth is machine-generated. Human-checked references would remove the benchmark’s central weakness.
- Nothing here generalises beyond archival print of this era, condition and script. Whether the pipeline works on other scripts is untested — an open Armenian Tesseract model exists and would make the cheapest first experiment.
Reproducing all of this
pip install -e '.[all]'
tetrak-ocr evaluate --all --save
Expect roughly 65 minutes, most of it Marker. Results land in
evaluation/ocr/benchmark.{csv,md} with a dated copy in evaluation/ocr/runs/.
Local engines reproduce exactly between runs; the Claude column does not, for the reason given above. Compare like with like by re-reading the dated CSVs rather than trusting a single pass.
Full detail lives in the reference section: the corpus with provenance and rights, the engines with each one’s constraints and failure cases, and routing with the decision table and bands.