Tetrak OCR
GitHub ↗
Reference

The engines

Seven engines behind one interface, plus two strategies built on them. Every one exposes ocr_image(path) -> str, so they are interchangeable; resolve one by name through the registry rather than importing it directly.

BackendExtra neededPDFsBest at
tesseractcore✓Baseline. Fast, offline, predictable
tesseract-autocore✓Same engine, per-image tuning. Best local average bar auto-local
vision[vision]✗macOS on-device engine. Strongest local result here, and the cheapest to run
easyocr[easyocr]✗Varied contrast and awkward grounds
paddle[paddle]✗Dense small text — but weak on this corpus
paddle-vl[paddle-vl]✓Document vision-language model. Strong, and slow enough to be batch-only
marker[marker]✓Layout-aware conversion; multi-column and PDFs
claude[claude]✓Anything the others cannot read
auto-localcore✓Unattended mixed material; runs several and keeps the best

See choosing a tool for which to use when, with the measurements behind it.

Tesseract, and auto-configuration

tesseract runs the engine with fixed settings. tesseract-auto calls analyse_image() first, which inspects the image and picks a contrast factor and page-segmentation mode per file.

That distinction matters more than it sounds. On the Carthay Circle postcard front, fixed settings score 0.00/0.00 and auto-configuration scores 0.49/0.86 — the same engine, on the same image.

EasyOCR and PaddleOCR

Both are deep-learning detector/recogniser pipelines, and neither reads PDFs — convert pages to images first.

On this corpus EasyOCR is middling (0.35/0.60) and PaddleOCR is weak overall (0.28/0.63) while winning two fixtures outright. That is a statement about this material rather than about PaddleOCR generally.

Both download model weights on first run, and hold them in memory afterwards, so the first call is slow and later ones are not.

Both join the auto-local pool for images, with EasyOCR ranked ahead of PaddleOCR on the scores above. EasyOCR was absent from that pool until recently: the original documentation described auto-local as using “paddle (or easyocr) for images”, but the code only ever offered PaddleOCR, so installing the EasyOCR extra bought nothing.

Adding it changed the measured outcome by nothing at all. EasyOCR wins the Carthay Circle front outright at 0.87, and still loses the selection there to a Tesseract transcript scoring 0.49, because the winner is ranked on quality weighted by word count. See results.

Armenian, with our own recogniser

easyocr-hy is stock EasyOCR with a different recogniser: tetrak_hy, trained for the Armenian script by tetrak-hy-trainer and published as tetrak-easyocr-armenian. Detection is untouched — CRAFT already finds Armenian text; reading it was the missing half. Stock EasyOCR scores 0.03 word recall on Armenian, which is a polite way of saying it reads nothing.

It has its own registry name rather than being folded into easyocr, for the same reason tesseract-auto does: the harness scores backends, so the difference between the two is a measurement instead of a claim.

Measured on eight evaluation registers — 65 proofread pages from seven works on Armenian Wikisource plus ten from volume 2 of the Armenian Soviet Encyclopedia, every page held out from training. The figures are means over the eight registers; the per-register tables are in the latest comparison:

BackendChar similarityWord recall
tetrak-hy v6 (ours)0.9600.908
hye-paddle0.7400.861
marker0.7820.803
hye-calfa-n0.9010.797
tesseract-hye0.8600.667

Means over the 8 evaluation registers, generated from evaluation/ocr/registers.csv at build time. Bold marks the best in each column. Only engines measured on every register appear.

The row labelled tetrak-hy v6 (ours) is easyocr-hy as it ships: the recogniser, its word list, the script fold and the layout below. The table names the recogniser rather than the backend, because the harness that emits it scores models. Stock EasyOCR is absent because it was only ever run on the encyclopedia pages (the 0.03 above), and the table admits only engines measured on every register — the rule that keeps a column’s mean over the same pages as the column beside it.

Since v6, easyocr-hy leads both means, and leads both metrics on most registers. It does not lead everywhere, and the exceptions have names:

Calfa’s two models are the strongest of the others. Both are CC BY-NC 4.0, so Tetrak can measure them but cannot ship or bundle either. easyocr-hy can be shipped commercially: its code and weights are Apache 2.0, and its word list is CC BY-SA 4.0, as the sources it is counted from require.

Three steps sit between the raw recogniser and those figures, and none is retraining:

Needs pip install "tetrak-ocr[armenian]". The recogniser’s weights (~15 MB) download on first use, pinned to one immutable revision per release and checksum-verified; EasyOCR’s detector weights come with them. No PDFs — convert pages to images first.

PaddleOCR-VL

A document vision-language model in the same paddleocr package as paddle, and a different thing entirely. Where paddle detects glyph boxes and classifies each one, paddle-vl generates the transcript conditioned on the image — a small vision encoder feeding an ERNIE decoder, Apache 2.0, about 0.9B parameters.

It is a separate backend rather than a flag on paddle, for the same reason tesseract-auto is separate from tesseract: the harness scores backends, so the difference between the two is a measurement instead of a claim.

Two things set it apart from every other local backend here.

It reads documents, not just images. PDFs and multi-frame TIFFs both work, returning one result per page with that page’s own text, so it is the only local backend that carries no multi-page refusal. marker and tesseract read PDFs too, but every other local backend — marker included — reads frame 0 of a multi-frame TIFF and calls reject_multi_page rather than return one page as though it were the document.

It is slow enough to change how you use it. On CPU a dense newspaper page took over ten minutes, against a fraction of a second for tesseract-auto. Treat it as a batch backend: excellent for a queue of scans overnight, wrong for anything interactive.

That cost is why it is opt-in to the auto-local pool rather than a member of it. auto-local runs every eligible backend on every file, so admitting this one by default would slow a routine call by two orders of magnitude — including on the postcards, where vision still beats it. Add it when the accuracy is worth the wait:

tetrak-ocr batch --backend auto-local --with-paddle-vl

It is worth the wait more often than not. On the corpus it averages 0.71 character similarity against auto-local’s 0.57, and it wins exactly where the existing pool is weakest: the two multi-column newspaper pages go from 0.28 and 0.39 to 0.78 and 0.92. It does not win everywhere — vision keeps the postcards and marker keeps the PDF — which is the argument for having both in the pool rather than swapping one for the other.

The pipeline version is pinned in the backend rather than left to the package default, because it moved twice in the month after we first looked at it, and an unpinned model means a routine pip install quietly re-measures the benchmark.

Inference is in-process only. The same class can talk to vLLM, SGLang and other servers, and takes a URL and an API key to do it; this backend names the native path explicitly and passes neither, so a local backend cannot start sending archive images to a third party because a default changed. A test asserts it.

Needs pip install "tetrak-ocr[paddle-vl]", which pulls paddlex[ocr] — not paddleocr[doc-parser], however much that sounds like the one. Model weights (~1 GB) download on first use and cache under ~/.paddlex/.

Marker

Layout-aware document conversion, and the only local backend that models reading order — which is why it wins the multi-column newsprint cover.

It is uneven: it also collapses on Grauman’s Chinese Theatre (0.09/0.00), where its layout model reads the whole postcard as a figure and returns nearly nothing. And it is slow: about 18 minutes of a 70-minute benchmark run on its own, and roughly 60 of those 70 once the two auto-local modes are counted, since both route to it.

Apple Vision

macOS only. pip install 'tetrak-ocr[vision]' — the extra pulls a PyObjC wrapper and nothing else, because Vision ships with the operating system. No weights, no download on first run, no network. Apple states that all of Vision’s processing happens on the device, so it sits inside the local-first guarantee exactly as Tesseract does.

On this corpus it is the strongest local engine measured, and the fastest worth using — which is why auto-local runs it second, immediately after Marker:

BackendCharacter similarityWord recall
vision0.440.78
marker0.420.62
tesseract-auto0.400.59
easyocr0.350.60
tesseract0.300.46
paddle0.280.63
auto-local0.520.71

Eight images, the PDF excluded because Vision cannot read one. Sub-second to four seconds per image, against Marker’s minutes.

The word recall is the striking number: 0.78 beats auto-local’s 0.71, which runs several engines and picks the best of them. Vision finds more of the text than anything else local.

Where it fails is reading order. On the Inside Facts newsprint it recovers 0.88–0.90 of the words and scores 0.08–0.12 on character similarity — the widest gap between the two metrics anywhere in this benchmark, and the same multi-column failure described in the research, in a more extreme form than Tesseract manages.

Two limits worth knowing:

No preprocessing is applied: Vision’s neural detector handles rotation and perspective itself, so the contrast and page-segmentation work behind tesseract-auto has no equivalent here.

Claude

The Anthropic vision API. Reads material no local tool can — it recovers 0.97 word recall on the vaudeville poster, where the best local backend manages 0.39.

Three caveats. It needs ANTHROPIC_API_KEY and it sends your images to an API, which rules it out for confidential material. Its benchmark scores are measured against ground truth it generated itself, so they are a ceiling reference rather than a measurement — see method. And unlike every local backend, it does not reproduce between runs: on the poster its character similarity has scored 0.73, 0.24 and 0.26 across three runs while word recall held at 0.96–0.98. It recovers the same words and orders them differently, which is what an order-sensitive metric does to a document with no single reading order.

Multi-page TIFF

Every backend accepts .tif/.tiff, and always did. What none of them handled was the multi-page case, which archival TIFF often is — a scanned pamphlet or register arrives as one file with a frame per leaf.

Pillow opens such a file at frame 0 and reports nothing about the rest, so “TIFF support” meant transcribing the cover and discarding the body. A partial transcript is the worst failure available to this pipeline: it passes quality scoring, enters the archive, and looks complete to everyone who reads it afterwards. The Claude backend already guards against the same shape of bug with its max_tokens check.

So a multi-frame image is now either read in full or refused by name:

BackendMulti-page TIFF
tesseract, tesseract-autoReads every page, joined by a blank line — the same way PDF pages already were
claude, easyocr, paddle, markerRaise MultiPageNotSupportedError, naming the page count and a backend that will read it
auto-localWorks: the frame-0 backends decline, Tesseract reads the whole file, and fan-out keeps that transcript

auto-local needed no special handling — it already treats a backend that raises as one that did not compete, so the complete transcript wins by default.

Not every extra frame is a page. TIFF’s NewSubfileType tag marks reduced-resolution renditions — pyramid levels and embedded thumbnails, which scanning pipelines emit routinely — and those are copies of a leaf already present, not further leaves. They are skipped, so a page plus its thumbnail counts as one page and is not refused by the frame-0 backends. Reading them would transcribe the same leaf twice: measured on the Carthay Circle postcard, appending a half-size level took the transcript from 32 words to 52.

Encoding costs more accuracy than paging does. CCITT G4 bilevel is the archival norm for text documents — tiny, and lossless for pure black-on-white. It is also where accuracy goes: the Hollywood playbill reads 297 words from its RGB original and 163 from a bilevel copy, a 45% loss, because thresholding to one bit destroys the antialiasing Tesseract uses to resolve small type. Converting back to RGB recovers nothing; the information is gone at source. If you control the scanning, keep greyscale.

Page settings are not adapted per frame. That is the conservative choice rather than a measured one: per-page auto-configuration lost badly on the PDF fixture, and there is no multi-page TIFF in the corpus to fit anything better against.

What it costs to run

Marker dominates the cost of any auto-local run, and the shape of that cost is not the one people expect.

Time tracks the amount of text recovered — not page count, not file size. On an M1 Max via Metal Performance Shaders, roughly 13–18 seconds per 1,000 characters of transcript:

DocumentPagesTranscriptTook
Trade weekly, dense multi-column10337,00090 min
Souvenir programme2133,0007 min
Illustrated book1012,0003 min
Playbill17,0005 min
Picture postcard117512 sec

A 21-page programme finished in 7 minutes; a 10-page trade weekly took 90. A single-page playbill took 5 minutes; a single-page postcard 12 seconds.

Recognition is autoregressive — decoded token by token, per detected region — so more text means proportionally more sequential work. That also rules out the obvious remedy: batch size does not help. Across the full range from 1 to CUDA-level values the spread is 6%, output is byte-identical, and the defaults are fastest. Surya’s MPS defaults being 4–8× below its CUDA ones looks like a handicap and is not one, because a larger batch cannot parallelise a decode loop.

For comparison on the same material, Tesseract runs at about 2.4 seconds per page. Naming a cheap backend is the way to trade quality for speed:

tetrak-ocr batch --backend tesseract-auto --quality-gate

The measurements behind all of this, including the experiments that came to nothing, are in design research note 001 (Marker throughput and multi-page PDFs).

auto-local

Not an engine — a strategy. It runs every viable local backend (Marker where a GPU makes it practical, then Vision on macOS, EasyOCR, PaddleOCR and auto-tuned Tesseract for images; Marker plus Tesseract for PDFs), scores each transcript with a reference-free quality score, and keeps the best. Best local result on this corpus at 0.57/0.76, and the slowest, since it runs several engines per file.

It needs the qa extra: scoring uses a spell checker and GPT-2 perplexity. See auto-local routing.

The quality gate is separate from the strategy

If the winning transcript still scores below the threshold, auto-local raises and the batch pipeline diverts that file to workspace/triage/ with a manifest.

The same gate is available to any single backend:

tetrak-ocr batch --backend tesseract-auto --quality-gate

It is opt-in rather than always on, because turning it on changes which files reach workspace/processed/ and nobody running --backend tesseract today should suddenly find output diverted. It needs the qa extra for the same reason auto-local does — the flag checks that up front and refuses, rather than failing partway through a batch.

auto-local’s own manifest is richer: it already scored every engine it ran, so it can show the whole table. A single backend produces one transcript and one number.

The retired fast mode

There was a second strategy, auto-local-fast, which ran only the top-ranked eligible backend. The benchmark retired it. In every installed configuration it was at best equal to naming a cheap backend directly, and on a GPU machine the ranking put Marker first — so “fast” mode resolved to the most expensive engine available and cost twenty-six times tesseract-auto’s runtime for less accuracy.

The one thing it uniquely offered was a single cheap engine with the quality gate. That is what --quality-gate is for. See results for the measurement that decided it.