Tetrak OCR
GitHub ↗
Evidence

The research in depth

Tetrak OCR exists because of a benchmark, not the other way round. This is that work end to end: why we ran it, what we measured, what it found, the finding we got wrong, and the four mechanisms the evidence forced into the product.

Read it in order and it is an argument. Jump around with the contents on the right and it is a reference.

This is the detailed version. The walkthrough covers the same ground in five beats, with the corpus and the results drawn out, and is the better place to start if you have not read either.

Why run a benchmark at all

A real archive is not a stack of uniform pages. The nine items in this corpus include a linen postcard, a chromolithograph poster, two pages of 1930 newsprint, a bilevel TIFF of a theatre playbill and a twenty-one-page souvenir programme. They differ in typography, contrast, layout and physical condition, and they were digitised by different people at different times.

The obvious way to build an OCR tool for that material is to pick a good engine, configure it well, and run it over the batch. Every vendor comparison, every “best OCR library” article, and most of our own instinct pointed that way.

We could not find evidence that it was the right shape. So before designing anything we built a fixed corpus, generated reference transcripts, and measured every engine we could install against every document. The result decided the architecture — and, in one case, unbuilt a feature that had already shipped.

What we measured

Nine items of Los Angeles stage and picture-palace ephemera, c. 1910–1945, chosen to span the failure modes that matter in archive digitisation rather than to flatter anything. Provenance and rights for every item are in the corpus reference.

Each transcript is scored on two numbers, both against normalised text — lowercased, whitespace collapsed.

Character similarity compares the transcript to the reference as a sequence. It is order-sensitive: text read in the wrong order scores badly even when every word is present.

Word recall is the share of expected words appearing anywhere in the transcript. It is order-insensitive, and it ignores extra words — an engine that invents text is not penalised.

They are always reported as a pair, because they fail differently and the gap between them is where the interesting behaviour lives:

PatternMeans
High recall, low similarityThe words were found and read in the wrong order — the signature of a multi-column page read as one column
High similarity, low recallWhat was read was read correctly, but some text was never seen — faint lettering, ornament, low contrast
Both lowThe engine could not read the page

On our deliberately unruly corpus, we found that no single tool performed well on all items, and gained valuable insight into where each tool stood out.

The weakness at the centre of the method

Reference transcripts were generated by Claude (claude-opus-4-8) and are committed under evaluation/ocr/corpus/expected/.

The reference is committed rather than regenerated per run precisely because of that instability; a moving reference would make run-to-run comparison meaningless. Human-checked ground truth would fix this properly, and it is the single highest-value improvement available to this work.

What the numbers support: relative comparison between local engines on this kind of material, and relative speed. The metrics are deterministic, the corpus is fixed and public, the reference is committed, so the local columns reproduce exactly.

What they do not support: any claim about Claude’s absolute accuracy, or generalisation beyond archival print of this era and condition.

The results

Character similarity / word recall, nine fixtures, run on an Apple MPS machine so auto-local had Marker available. Bold marks the best local engine per fixture; N/A means the engine cannot read PDFs.

Fixturetesseracttesseract-autoclaude*easyocreasyocr-hypaddlepaddle-vlmarkervisionauto-local
bancroft-magician-poster.jpg0.07/0.000.28/0.000.72/0.770.40/0.140.17/0.000.55/0.320.31/0.140.17/0.000.57/0.410.57/0.41
carthay-circle-postcard-back.png0.59/0.500.93/0.790.98/0.920.79/0.620.49/0.080.28/0.620.68/0.790.66/0.620.65/0.830.93/0.79
carthay-circle-premiere.jpg0.00/0.000.51/0.860.89/0.930.87/0.710.74/0.290.09/0.000.88/0.860.76/0.860.89/0.930.51/0.86
chinese-theatre-triptych-postcard.jpg0.44/0.400.44/0.401.00/1.000.50/0.670.46/0.000.79/0.730.83/0.870.45/0.470.66/0.930.45/0.47
graumans-chinese-theatre.jpg0.66/0.440.79/0.440.88/0.560.72/0.500.61/0.060.74/0.190.72/0.310.10/0.000.73/0.380.73/0.38
greek-theatre-night.jpg0.90/0.750.90/0.751.00/1.000.95/0.750.79/0.000.91/0.500.99/0.920.74/1.000.98/1.000.74/1.00
hollywood-boulevard-east.jpg0.63/0.330.62/0.330.93/0.830.72/0.440.65/0.170.64/0.440.65/0.390.11/0.000.68/0.390.68/0.39
hollywood-boulevard-west.jpg0.42/0.270.42/0.270.96/0.830.57/0.300.42/0.170.58/0.430.49/0.270.08/0.000.53/0.430.42/0.27
hollywood-music-box-playbill-1926.tif0.52/0.460.57/0.530.99/0.960.50/0.390.34/0.260.78/0.910.88/0.900.47/0.740.69/0.810.47/0.74
inside-facts-1930-cover.jpg0.68/0.850.90/0.911.00/0.990.41/0.710.65/0.500.50/0.940.98/0.930.75/0.900.63/0.880.75/0.90
inside-facts-1930-page-six.jpg0.94/0.910.94/0.910.99/0.950.30/0.780.61/0.600.30/0.970.96/0.950.74/0.930.59/0.890.74/0.93
kar-mi-troupe-poster.jpg0.01/0.000.09/0.010.99/0.980.12/0.390.10/0.000.27/0.380.29/0.440.03/0.000.29/0.460.29/0.46
kinema-theater-ad-1920.tif0.45/0.400.45/0.400.75/0.980.42/0.670.43/0.330.49/0.930.74/0.920.93/0.870.88/0.910.93/0.87
king-of-kings-souvenir-1927.pdf0.91/0.940.91/0.940.93/0.97N/AN/AN/A0.89/0.900.87/0.92N/A0.87/0.92
over-the-fence-poster.jpg0.19/0.140.40/0.180.77/0.820.49/0.230.27/0.140.64/0.450.39/0.360.19/0.000.54/0.450.64/0.45
parlor-match-poster.jpg0.00/0.000.14/0.031.00/1.000.67/0.440.24/0.000.68/0.310.88/0.720.76/0.750.79/0.660.79/0.66
thurston-magician-poster.jpg0.08/0.040.26/0.260.99/0.960.74/0.520.12/0.000.74/0.440.73/0.630.72/0.520.62/0.520.72/0.52
Average0.44/0.380.56/0.470.93/0.910.57/0.520.44/0.160.56/0.540.72/0.660.50/0.500.67/0.680.66/0.65

Generated from evaluation/benchmark.csv at build time. Bold marks the best local backend per fixture. claude generated the ground truth and is a ceiling reference rather than a score.

Fixturetesseracttesseract-autoclaude*easyocreasyocr-hypaddlepaddle-vlmarkervisionauto-local
bancroft-magician-poster.jpg0.50.33.93.04.010.923.18.90.915.1
carthay-circle-postcard-back.png0.40.56.20.40.47.221.25.70.616.8
carthay-circle-premiere.jpg0.20.52.40.30.36.818.15.60.515.3
chinese-theatre-triptych-postcard.jpg0.30.32.20.30.36.013.74.40.513.2
graumans-chinese-theatre.jpg0.30.32.40.30.36.316.65.40.513.5
greek-theatre-night.jpg0.30.32.40.30.34.315.15.40.512.3
hollywood-boulevard-east.jpg0.30.32.80.30.37.411.05.60.515.5
hollywood-boulevard-west.jpg0.30.33.50.30.36.417.96.80.514.8
hollywood-music-box-playbill-1926.tif1.01.128.15.77.842.6170.3274.71.2325.3
inside-facts-1930-cover.jpg2.22.227.27.89.757.5221.1286.82.2349.6
inside-facts-1930-page-six.jpg4.94.964.017.721.5132.4612.5308.04.2461.4
kar-mi-troupe-poster.jpg0.30.411.20.71.09.227.02.90.816.8
kinema-theater-ad-1920.tif0.80.817.32.73.826.3126.523.21.158.3
king-of-kings-souvenir-1927.pdf35.535.2150.7———1128.7423.3—431.6
over-the-fence-poster.jpg0.30.33.00.30.36.514.73.20.612.9
parlor-match-poster.jpg0.40.33.50.40.67.925.77.00.819.3
thurston-magician-poster.jpg0.30.33.20.30.46.318.16.40.619.4
Total seconds48.248.3333.940.851.4344.02481.41383.316.11811.0

Generated from evaluation/benchmark.csv at build time. claude generated the ground truth and is a ceiling reference rather than a score.

Average character similarity and word recall for each backend, with Claude separated as a ceiling reference Average character similarity and word recall for each backend, with Claude separated as a ceiling reference

Both the table and the chart are generated from evaluation/ocr/benchmark.csv at build time, so neither can drift from the data.

What the numbers say

No engine wins across the board

This is the headline, and it is why the average column is the least useful part of the table. Five different engines win across the nine fixtures:

An earlier version of this page added “and none wins more than twice”. That is no longer true — Marker takes three, both TIFFs having landed in territory it handles well. The claim that matters is unchanged: no engine wins most of them, and the winner is not predictable from the average.

Route by document, not by average.

The metrics disagree, and that is the point

On newsprint every local engine pairs high word recall (0.71–0.94) with near-zero character similarity (0.03–0.45). The words are all there; the columns are interleaved. Character similarity is order-sensitive and word recall is not, so the gap between them is the layout failure, quantified.

PaddleOCR recovers 94% of the words on the Inside Facts front page and scores 0.21 on similarity. For a human reader that output is unusable; for full-text search it is adequate. Which number matters depends on whether you are building a reading edition or a search index, and that is your question rather than ours — which is why we never collapse the two into one score.

Tuning beat switching

tesseract-auto averages 0.41/0.63, ahead of EasyOCR at 0.35/0.60 — a deep-learning engine beaten by Tesseract with better per-image settings.

It has not always been. It previously scored 0.22 and lost to plain Tesseract. The fix was re-fitting the page-segmentation bands in analyse_image(), which had been inherited from a parent project and were sending dense pages to the sparse-text mode. Carthay Circle’s front went 0.00/0.00 → 0.49/0.86 on that change alone.

The same engine, the same image, a different page-segmentation mode. Reaching for a heavier engine is the expensive reflex — gigabytes of weights, minutes of runtime, sometimes an API bill — and on this evidence it is often the wrong first move.

…but the first tuning did not generalise

Two TIFFs were added after analyse_image()’s thresholds were first fitted by hand, which made them the only held-out data here. Auto-configuration was worth nothing on either — it gained nothing on the playbill and lost 0.19 on the Kinema ad, against a gain of +0.15 on the seven it had been fitted to. On the fitted set it gained on five fixtures, once by +0.49. That is what overfitting looks like, and with seven images against four contrast bands and three PSM bands there was ample freedom for it to happen.

The Kinema failure was diagnosable rather than mysterious. Its frame is 42.9% dark because a large illustration fills half of it, which put it in the “image-dominated, no page structure worth analysing” band and selected PSM 6, a single uniform block. But it does have structure: a narrow programme column in tiny type beside the headline block. PSM 3 read it at 0.23; PSM 6 collapsed it to 0.04.

analyse_image()’s own docstring anticipated the risk — “it cannot tell ink from imagery” — and reasoned that a dark frame has little structure to analyse anyway. That held for the night-scene postcards the band was fitted against. It does not hold for an illustrated page that still has columns, and a histogram cannot tell those two apart.

Re-fitting them, and what it cost

On 27 September 2026 the bands stopped being hand-written. They are now chosen by evaluation/ocr/calibration/fit.py, which scores all 42 (contrast, psm) pairs against every raster fixture and picks the band edges that minimise mean regret — how far each fixture falls below the best it could have scored. Regret rather than mean score, because the corpus spans a 350× range in transcript length and a mean hands the objective to the longest documents. The current values are on the routing reference, read from the module rather than typed into the page.

Five of the nine tesseract-auto figures moved, and all five moved up. The Kinema ad now matches plain Tesseract exactly instead of collapsing; the Inside Facts inner page, which had been quietly losing 0.09 to the contrast band, gains it back; the Carthay Circle premiere, the playbill and the Kar-Mi poster each gain a little. The postcard reverse, Grauman’s and the Inside Facts cover were already on their fitted cell and did not move, and the PDF cannot — the bands never run on that path. Nothing scores below plain anywhere in the corpus now, where two fixtures did.

Repairing the Kinema ad meant spending the held-out data. The PSM-6 edge was unconstrained anywhere between 25.3% and 56.1% dark on the six fitted fixtures, and the Kinema ad sits at 42.9% — inside that gap. No fit that could not see it would ever have moved the edge off it. So it was promoted into the fit set, and for a day there was nothing held out.

That left leave-one-out as the only generalisation measure, and it was not reassuring: refitting without each fixture in turn and scoring it with the bands that result gave a mean regret of 0.20 against 0.04 in sample. The conclusion drawn at the time was that eight fixtures was the binding constraint, and that more scans would fix it.

Doubling the corpus, and what it actually showed

On 28 September the corpus went from nine fixtures to seventeen. Eight rasters — four linen postcards and four lithographed theatre posters — had been sourced in August specifically to widen calibration and had sat unusable ever since, because none of them had a reference transcript. Generating those, and then checking every transcript in the corpus by hand for the first time, restored a genuinely held-out set of four.

The check was worth more than the widening. Five of the seventeen references were wrong. Two omitted a printer’s imprint entirely; one read a fragment of musical notation as the number 18; two were outright fabrications on the Kar-Mi poster, where “drinking water through a tube” had become “drinking water successfully”. And one was worth 0.25 character similarity on its own: the Inside Facts inner page had been reflowed into continuous paragraphs, joining 27 end-of-line hyphenations that the newspaper actually prints. Every engine had been penalised for reading the page correctly.

A second Kinema turned up immediately, in material nothing had been fitted to. The Greek Theatre postcard is 69.4% dark — a night scene — so the bands read it as image-dominated and sent it to PSM 6, scoring 0.37. Plain Tesseract reads it at 0.90. The dark_pct heuristic fails at the top of its range as well as in the gap the Kinema ad sat in, which the original corpus could not have shown because the Kinema ad was the only dark-framed fixture in it.

And generalisation got worse, not better. Leave-one-out regret is 0.265 over twelve fitted fixtures, against 0.20 over eight. More data made the problem harder, because the material is now genuinely varied: postcards, newsprint, a playbill and chromolithographs do not share one mapping from two pixel statistics to a configuration. Nor does spending more bands help — the figure is identical at four contrast bands and at five. Neither the corpus size nor the model’s complexity is the limit. Two features are.

That is a more useful answer than the one the smaller corpus gave, and it is the opposite of it.

Note what did not happen when the two TIFFs were originally added: no other engine’s ranking changed, and auto-local still routed the Kinema ad correctly by picking Marker at 0.75. The quality score caught what the heuristic got wrong — which is the argument for measuring transcripts rather than predicting from pixels.

Fan-out wins, by less than it should

auto-local averaged 0.53/0.74 before 28 August 2026, the best local result on both metrics, and its margin over tesseract-auto widened from 0.05 to 0.12 as the corpus grew, because fan-out measures each transcript instead of predicting from pixels. It now averages 0.57/0.76 — Vision joined the candidate pool that day too, alongside the fix below, so the margin over tesseract-auto is now 0.16.

The ceiling is higher, though closer than it was. An oracle taking the best local engine on every fixture averages 0.61; auto-local reaches 0.57, about 94% of it. Four fixtures accounted for the gap before the fix below; two of them were a scoring problem, not an engine one, and are fixed:

FixtureBest localChose beforeChose afterCost before → after
carthay-circle-premiere.jpgvision 0.89tesseract-auto 0.49marker 0.760.40 → 0.14
graumans-chinese-theatre.jpgtesseract-auto 0.86paddle 0.78vision 0.830.08 → 0.03
hollywood-music-box-playbill-1926.tifpaddle 0.39marker 0.28marker 0.280.11 (held out; not fit to)
inside-facts-1930-page-six.jpgtesseract 0.45marker 0.40marker 0.400.05 (plain tesseract is not a candidate)

The last row does not move, and cannot: auto-local never runs plain tesseract, only tesseract-auto, so no scoring change closes a gap against a backend the router does not have. The third is held out from fitting on the same principle as the Tesseract auto-configuration bands above — “…but the tuning does not generalise”.

The Carthay Circle case repays reading closely, because it shows the raw quality score getting a candidate genuinely wrong, not just outranked by length. Marker’s combined_score (0.1517) is fractionally the highest of the four candidates — ahead of Vision’s (0.1405) and EasyOCR’s (0.1416) — even though Marker’s character similarity (0.76) trails both. Under the old ranking, quality × (0.5 + 0.5 × words/max_words), none of that mattered: tesseract-auto emitted 77 words (mostly noise from the postcard’s textured background) against 11 each for the other three, and word count alone overrode all of them, promoting a transcript scoring 0.49 over one scoring 0.89 — the lowest-scoring candidate winning outright.

Raising the weight to 0.8 removes the length effect for this fixture entirely — tesseract-auto’s own combined_score is the lowest of the four, so no plausible weight would still pick it — which is enough to hand the result to Marker instead. It cannot fix Marker’s ranking error, because that error was never about length: evaluation/ocr/calibration/router_sweep.py fits the weight by minimising mean regret against ground truth, the same principle the Tesseract bands are held to. Under the corrected character similarity (October 2026), mean regret over the fit set falls from 0.134 to 0.090 and 0.8 still wins, though one fixture, graumans-chinese-theatre, now scores fractionally worse; under the metric in use when the weight was chosen, the fall was 0.083 to 0.032 with no fixture worse. the error taxonomy (product analytics) and brief 010 (the router length weight) have the full account, including why the residual 0.14 is now a different, smaller problem, inside combined_score rather than the router.

A feature we measured and then removed

auto-local-fast ran only the top-ranked eligible engine instead of all of them. On this machine that was always Marker, so it produced identical scores to marker (0.42/0.61) at almost identical cost — strictly worse than an engine already available:

character similarityseconds
tesseract-auto0.4945
auto-local-fast0.421169

Lower accuracy and twenty-six times the runtime. Nor was that only a GPU-machine problem: worked through every installed configuration, it never won. tesseract-auto already occupied the niche fast mode was invented for and filled it better, so the backend was removed, one day after it shipped.

Registering it as a backend rather than hiding it behind a flag is what turned a preference into a measurement, and then into a deletion.

Marker is uneven and slow

It wins the newsprint cover, the Kinema ad and the PDF, then collapses on Grauman’s (0.09/0.00), where its layout model reads the whole postcard as a figure and returns nearly nothing.

It also dominates the runtime. Of the roughly 65 minutes a full --all run takes, about 50 are Marker. Dropping it from the fan-out pool would cost more than it saves: it wins three of nine, including the Kinema ad at 0.75 where the next best local result is plain Tesseract at 0.23.

The poster defeats everything local — and destabilises the reference

Best local character similarity is 0.13; retuned Tesseract manages 0.02.

Claude is the only thing that reads it, but how well is not a stable number. Across four runs against the same committed ground truth — which Claude itself generated — it scored 0.73, 0.24, 0.26 and 0.73. Word recall over the same four runs was 0.98, 0.96, 0.97 and 0.98.

That split is the finding. Claude recovers the same words every time and arranges them differently every time, because a chromolithograph has no single correct reading order for a transcript to be scored against. It is the same signal as the interleaved newsprint columns, arriving for a different reason.

Hand-lettered display type is where matching letter shapes stops working and understanding what the page is starts mattering. If your material looks like this, the local tools will not do — and this benchmark cannot tell you precisely how much better the alternative is.

A finding we retracted

An earlier version of this analysis claimed auto-local mis-ranked the PDF, choosing Marker when Marker had scored 0.02 character similarity there.

That was wrong and has been retracted. The 0.02 was an artefact of truncated ground truth — the Claude backend was calling the API with max_tokens=4096 and silently returning partial transcripts for the twenty-one-page programme, so Marker’s correct output was being compared against a fragment. Against the corrected reference Marker scores 0.66 and is the best engine on that fixture. The quality heuristic had been right all along.

It is recorded rather than quietly deleted because the failure mode is instructive: a truncation bug in the reference looked exactly like a routing bug in the system under test. The Claude backend now raises on truncation rather than returning partial text.

What the evidence built

Four findings, four mechanisms. Together they are auto-local, and together they are the six stages between dropping a file in and getting a transcript out.

Six stages of the pipeline, drawn as an archivist's workroom: an ingestion desk taking TIFFs, JPEGs, PNGs and PDFs into the scans directory; optical calibration, where analyse_image() reads pixel intensity and dark ratio to set the contrast and page-segmentation mode; sequential evaluation, where each engine runs in turn rather than in parallel to avoid contention for CPU and memory; a lexical jury drawn as two dials, one reading real words and the other reads like language, both of which must register; a triage desk where anything scoring below 0.10 is held for a person to decide, for manual review or escalation to Claude; and the finished archive, where clean Markdown is filed beside the original scan.
What happens between dropping a file in and getting a transcript out. The two dials in panel four are the quality score, and the reason there are two: dictionary coverage catches real letters that are not real words, while perplexity catches real words in an order that means nothing. The score multiplies them, so neither rescues the other.

Fan out, and select per file

No engine wins, so the obvious design — pick one, configure it, apply it to the batch — is unavailable. Any single choice is wrong for most of the corpus.

auto-local runs every viable engine over each file and keeps the best transcript. Selection happens per file, not per batch, which is the only granularity at which the finding can be acted on. The cost is honest: several engines per file, so it is the slowest local mode.

The engines run sequentially, deliberately. Each is heavy on CPU and memory and several hold model singletons; starting two or three at once trades predictable runtime for contention, and it hurts most on exactly the large batches where the saving would matter.

Auto-configure before reaching for a bigger engine

Tuning beat switching, so analyse_image() inspects each image before OCR and picks contrast and page-segmentation mode from two cheap statistics:

SignalWhat it tells usEffect
Standard deviation of pixel intensityHow tonally flat the scan is — faded ink on aged paper reads lowContrast boost
Share of frame below luminance 128How much of the frame is not paperPage-segmentation mode

The second misleads by its name. dark_pct looks like a text-density measure and is not: a night-scene postcard reads 56% dark because of the photograph. That turned out to be useful anyway — a frame dominated by imagery has little page structure for layout analysis to work with, however the darkness got there. The heuristic works for a slightly different reason than the one it was reached for, which is also why the hand-fitted bands failed on the Kinema ad above: the signal is sound, and the edges drawn across it were not.

The exact bands are in the routing reference, read from the module at build time rather than written out here — they are fitted from a sweep and they move.

Score without a reference

The metrics disagree, and both need ground truth. At run time there is none, so whatever picks the winner has to judge a transcript it has never seen the answer to.

Two signals that fail differently, multiplied, then weighted by relative word count:

The word-count weighting came straight out of an observed failure: on sparse images a conservative engine emitting eight clean words was beating one that recovered twenty-three noisier ones, because dictionary coverage is trivially 1.00 for a short clean phrase. Quality alone rewarded saying less. As the Carthay Circle case shows, that correction now overshoots in the other direction on at least one fixture.

Fail loudly

Some pages defeat everything local. For a tool running unattended over an archive that is the dangerous case — not because it fails, but because a failed transcript looks exactly like a successful one from outside. A .md file appears in workspace/processed/, and nothing signals that it is noise.

So there is a quality floor of 0.10. Below it the scan is diverted to workspace/triage/ with a manifest recording what every engine produced and how each scored. It never reaches workspace/processed/.

The point is not that the pipeline handles hard documents. It is that you can tell which ones it did not, which makes “what is still in triage/” a meaningful question and an actionable queue.

What we chose not to build

No automatic escalation. The triage queue is passive. Forwarding failures to a vision model automatically would be easy, and would turn a local, free, offline tool into one making paid API calls on material the user may not have wanted to leave the machine. That decision stays with a person.

No single quality number. Character similarity and word recall are reported as a pair everywhere, because collapsing them hides the finding above.

No claim about Claude’s accuracy. It generated the ground truth it is scored against. It sets a ceiling reference and nothing more.

Choosing a tool for your own material

The practical version of “no engine wins”:

Your materialUseSecond choiceWhy
Clean printed captions, postcards, labelstesseract-autoeasyocr0.91 against Claude’s 0.98. Free, offline, fast
Multi-column newsprint, magazinesclaudemarkerReading order is the whole problem, and only layout understanding solves it
Decorative, hand-lettered, ornamental typeclaude—Nothing local reads it. Best local score is 0.13
Multi-page PDFsmarkertesseractEasyOCR and PaddleOCR cannot read PDFs at all
Mixed material, unattendedauto-local—Runs the viable local engines, keeps the best transcript
Anything confidentialauto-localtesseract-autoEverything stays on the machine; no API call

On the Carthay Circle postcard reverse, five points of character similarity separate a free offline tool (tesseract-auto, 0.91) from a paid API call (Claude, 0.96). For clean printed text that is the norm rather than the exception, and it is the case where reaching for a vision model is hard to justify.

If you do route decorative and multi-column material to Claude, a page runs roughly 1,500 input and 750 output tokens on this corpus:

ModelInputOutputApprox. per page
Haiku 4.5$1 / MTok$5 / MTok~$0.005
Opus 4.8$5 / MTok$25 / MTok~$0.026
Opus 4.8, Batch API50% off50% off~$0.013

A practical pattern: run auto-local over everything and send only what lands in workspace/triage/ to Claude. On this corpus that would be the poster and little else.

What is still open

Reproducing all of this

pip install -e '.[all]'
tetrak-ocr evaluate --all --save

Expect roughly 65 minutes, most of it Marker. Results land in evaluation/ocr/benchmark.{csv,md} with a dated copy in evaluation/ocr/runs/.

Local engines reproduce exactly between runs; the Claude column does not, for the reason given above. Compare like with like by re-reading the dated CSVs rather than trusting a single pass.

Full detail lives in the reference section: the corpus with provenance and rights, the engines with each one’s constraints and failure cases, and routing with the decision table and bands.