Tetrak OCR
GitHub ↗
Reference

Auto-local routing

One strategy over the whole set of local engines, for when you would rather not choose a backend per document.

auto-local runs every viable local engine on a file, scores each transcript with a reference-free quality score, and keeps the best. It is the best local option on this corpus (0.57/0.76), and the most expensive, since it runs several engines per file.

The decision table below is implemented in tetrak_ocr.auto_local and asserted, case by case, in tests/test_auto_local.py. Every eligible candidate runs; the order settles ties.

One candidate is not in the table because it is not in the pool unless you ask. --with-paddle-vl adds PaddleOCR-VL to every branch below, images and PDFs alike, and ranks it first. It measures as the strongest local engine on the corpus and wins the registers this pool reads worst, but it runs in minutes where the rest run in seconds — and fan-out pays that cost on every file, so it is a choice rather than a default.

Decision logic for tetrak_ocr.auto_local, tetrak_ocr.qa_score, and the triage queue in tetrak_ocr.batch.

Yes — and Marker installed

No

Yes

No — image

Yes

No — image

No — passes quality gate

Yes — LowQualityError

runs via

Tesseract-auto detail (tesseract.py — analyse_image)

Yes

No

fitted band

fitted band

PDF?

Extract pages via pdf2image
Defaults: contrast 2.0 · PSM 3
auto-analysis skipped for PDFs

Convert to greyscale

StdDev of pixel intensities
(tonal range proxy)

Contrast factor
looked up in the fitted stddev bands

Dark pixel %
pixels with luminance < 128

Page segmentation mode
looked up in the fitted dark_pct bands

Tesseract OCR

Triage queue (batch.py — _send_to_triage)

Write workspace/triage/‹stem›.md
per-backend score table
+ first 500 chars of each transcript

Move original file to workspace/triage/‹filename›

Score each transcript (qa_score.py)

dict_coverage
fraction of word tokens found in an English
dictionary — catches non-word OCR noise
(pyspellchecker, ~100 KB, no model load)

combined_score = dict_coverage × 1 / log1p(perplexity)

perplexity
GPT-2 coherence score — lower means more
natural English; catches real-word substitutions
that spell-checking misses (transformers, ~500 MB)

effective_score = combined_score × (0.8 + 0.2 × words / max_words)

Word-count factor prevents a conservative backend
that emits few clean words from beating a more
complete one on sparse-text images

word_count per transcript
max_words = max across all candidates

Fan-out — run every candidate, collect transcripts

Marker

Vision

EasyOCR

PaddleOCR

Tesseract-auto

Input file

has_gpu()
CUDA · ROCm · MPS

PDF?

PDF?

Marker · Tesseract-auto

Marker · Vision¹ · EasyOCR¹ · PaddleOCR¹ · Tesseract-auto

Tesseract-auto

Vision¹ · EasyOCR¹ · PaddleOCR¹ · Tesseract-auto

Select winner
= candidate with highest effective_score

winner's raw
combined_score
< 0.10?

Return best transcript

workspace/triage/
awaiting manual review
or escalation to Claude

¹ Vision, EasyOCR and PaddleOCR are excluded from PDF candidates — none of them reads PDFs. Vision is additionally macOS-only, so off macOS it never appears as a candidate at all.

Candidates are ordered by measured result, because fast mode takes the first and fan-out uses the order to break ties. EasyOCR sits above PaddleOCR: it is ahead on every image where either engine reads anything at all. See results for the scores.


Notes

GPU detection (has_gpu())

Uses PyTorch, which is already installed as a transitive dependency:

This verifies the GPU is usable by the ML stack, not just visible to the OS.

Fan-out vs routing

The old auto-local picked one backend per file type before processing began. The new version runs every eligible backend and scores each transcript. The cost is roughly proportional to the number of backends × per-image time, which is acceptable given the priority of output quality over build speed.

Effective score: quality × relative word count

Pure combined_score was insufficient on sparse-text images — a conservative backend emitting 8 perfectly clean words can outscore a backend that recovers 23 slightly noisier words, because dict_coverage is trivially 1.00 and perplexity is low for a short well-formed sentence. The 0.8 + 0.2 × words/max_words multiplier ensures coverage is weighted alongside quality. The floor of 0.8 means even a zero-word backend retains most of its raw quality score, and max_words is relative so the formula is scale-invariant across different document sizes.

The weight was 0.5 until 28 August 2026. At that value a noisy 77-word tesseract-auto transcript on carthay-circle-premiere.jpg beat 11-word transcripts from three other engines on word count alone, despite scoring lowest of the four on raw quality — see the results. evaluation/ocr/calibration/router_sweep.py re-fitted the weight by minimising mean regret against ground truth; LENGTH_WEIGHT’s provenance comment in auto_local.py has the full result.

Quality gate threshold

MIN_QUALITY_THRESHOLD = 0.10 in qa_score.py. The gate uses the winner’s raw combined_score, not the effective score — the triage decision should reflect output quality, not quality × length.

Triage queue

Triage is passive by design: nothing escalates automatically. A file lands there with a manifest, and a human decides whether to send it to the Claude backend, re-scan it, or transcribe it by hand. Auto-escalation would quietly turn a local, free, offline pipeline into one that makes paid API calls on material the user may not want leaving the machine.

Tesseract auto-tuning

The diagram above shows the shape of the decision and deliberately carries neither the band edges nor the settings they select, because both move: they are fitted from a measured configuration sweep by evaluation/ocr/calibration/fit.py, and the first honest re-fit changed every edge. It would also be easy to read an ordering into the diagram that is not there — the flattest scans do not get the strongest contrast boost, which was the original premise and did not survive measurement. The current values, read from tetrak_ocr.backends.tuning when this page is built:

RangeSetting
0 <= stddev < 34contrast 2.0
34 <= stddev < 40contrast 2.5
40 <= stddev < 46contrast 2.0
46 <= stddev < 67contrast 3.0
stddev >= 67contrast 2.0
0 <= dark_pct < 10psm 12
10 <= dark_pct < 49psm 3
49 <= dark_pct < 69.2psm 6
dark_pct >= 69.2psm 3

Fitted to wide-corpus, from evaluation/ocr/corpus/splits.toml. Mean character similarity 0.5796; leave-one-out regret 0.2650 against 0.0760 in-sample.

The three segmentation modes the psm bands choose between: 3 is fully automatic page segmentation, which finds and orders columns; 6 assumes one uniform block, and is right when there is no real layout; 12 is sparse text with orientation detection, and its OSD is what reads the 90-degree rotated imprint on a postcard reverse.

PDFs bypass per-image analysis — pdf2image renders each page before any analysis step, so fixed defaults (contrast 2.0, PSM 3) are used; the bands never run on the PDF path, which is also why no PDF is fitted to. Note that dark_pct measures how much of the frame is not paper, not text density — a night-scene postcard reads 56% dark because of the photograph. That turns out to be the useful signal anyway: a frame dominated by imagery has little page structure for layout analysis to work with, however the darkness got there.