Tetrak OCR
GitHub ↗
Benchmark

Beating the baselines, starting with our own ruler

We set out to lead every Armenian OCR engine we can measure on every register. Before training anything, we found that most of the gap we were chasing was in our own metric. The rest came down to reading order, a word list and a second fine-tune: v6 now leads both means and most registers outright

At the end of September, the comparison with every external engine left us level on word recall and well behind on character similarity. Calfa’s Tesseract model led the order-sensitive metric on almost every register. We wrote that what survived reading order was a recognition-sharpness gap, and that it was the next thing to train against.

That turned out to be mostly wrong, and the way it was wrong is the more useful half of this article.

The ruler was bent

Our character similarity is difflib.SequenceMatcher(...).ratio(). SequenceMatcher has a parameter, autojunk, that is on by default. On any string longer than 200 characters, it treats every character making up more than one per cent of the text as junk and stops matching on it. A page of Armenian is a few thousand characters, and almost the whole alphabet, plus the space, clears one per cent. So the metric was matching on a handful of rare characters, and its score depended on where those few blocks happened to fall, not on how much of the page had been read.

The effect was not small. Single-column literary pages that we were reading almost perfectly had been scored as though about a quarter of the text were wrong. The trainer’s own alignment code had turned the flag off long ago, for exactly this reason. The metric module never had.

Fixing it meant re-running every external engine on every page, because we had kept their scores but not their output. They now save their raw readings, so the next correction to the metric is a re-score that takes seconds.

Three smaller corrections followed, each of the same kind: a transcription habit the metric was charging to whichever engine did not share it.

None of these changed which engine led which register on its own. Together with autojunk, they changed what the gap was.

What the gap actually was

With the metric fixed, the single-column registers were close: a hair behind the leader on both metrics, where they were behind at all. Two things were not.

Reading order on the encyclopedia. The Armenian Soviet Encyclopedia is set in three columns, closer together than EasyOCR pads its text boxes. No vertical line on the page was free of boxes. Our XY-cut found a gutter, but treated every box that overhung it as a full-width divider, such as a running head. It peeled the page apart a line at a time until its recursion limit gave up and read the rest straight across. The fix:

Without touching the weights, the encyclopedia went from our worst register on character similarity to one where we lead every engine we measure.

Recognition on three registers. Faustus of Byzantium, Tumanyan’s academic edition and the encyclopedia still trailed on word recall, by a little.

Closing it

Three changes, in order of what each was worth.

A word list. The recognition head has no language model, so a word it misreads by one letter comes out as written, even when the right word was its own second choice. We built a list of Armenian words from proofread Wikisource transcripts and the Nayiri Armenian Lexicon, with the evaluation pages excluded. For each word whose reading is not in that list, a small beam search over the recogniser’s own probabilities looks for a listed reading nearly as probable. We chose the threshold on training crops held out from the fine-tune, not on the evaluation pages. On those pages it fixed many times more words than it broke, and those it broke are proper nouns, classical spellings and edition spellings: the words any lexicon breaks.

A second fine-tune. v6 is v5 trained on twice as many real crops, cut from more volumes of the editions we were losing on. Two things went into the training data deliberately: the Armenian full stop in place of the transcribers' colon, and lines set in capitals for the encyclopedia’s headwords. Before training, a check of every harvested page against every evaluation page found 25 pages that print an evaluation page’s text: variants and reprints in an academic edition. One of them shared two hundred eight-word runs with the evaluation page it reprints. Those pages had been in v5’s training text. They are excluded from everything new in v6, but v6 starts from v5’s weights and reuses v5’s synthetic crops, so the exposure is inherited, not removed: the model card records it as a known defect on the four registers concerned (Tumanyan, Baronian, Faustus and the encyclopedia). Removing it takes a pre-train from scratch.

A digit fold. The Faustus index is set in a bold italic in which the model reads the digit 2 as the Armenian letters Չ or շ. Inside a number, neither letter can be meant, so the script fold now reads them as the digit.

Registertetrak-hy v6 (ours)hye-calfa-nhye-paddlemarkertesseract-hye
Totovents, Works0.9680.9410.9410.8070.739
Medical encyclopedia0.9680.8290.9330.9140.877
Otyan, Works0.9650.9290.9340.8370.727
Baronian, Works0.9440.9110.9180.7870.744
Faustus of Byzantium (1968)0.9160.9140.9220.8600.716
Tumanyan (academic edition)0.9140.9040.9130.7080.710
Encyclopedia (ASE vol. 2)0.9020.8430.8650.8360.773
Dictionary (hy–en)0.6880.1090.4570.6780.046
Mean over registers0.9080.7970.8610.8030.667

Word recall — order-insensitive, so it measures recognition almost independently of layout. Generated from evaluation/ocr/registers.csv at build time. Bold marks the best engine per register. Only engines measured on every register appear.

Registertetrak-hy v6 (ours)hye-calfa-nhye-paddlemarkertesseract-hye
Totovents, Works0.9870.9840.9850.8140.938
Medical encyclopedia0.9650.9750.3020.8880.973
Otyan, Works0.9870.9780.9810.9200.923
Baronian, Works0.9860.9820.9830.8640.943
Faustus of Byzantium (1968)0.9700.9760.9270.9190.929
Tumanyan (academic edition)0.9560.9500.9700.6580.889
Encyclopedia (ASE vol. 2)0.9700.9600.3340.5260.946
Dictionary (hy–en)0.8570.4020.4350.6630.339
Mean over registers0.9600.9010.7400.7820.860

Character similarity — order-sensitive, so it collapses where the reading order is wrong. Generated from evaluation/ocr/registers.csv at build time. Bold marks the best engine per register. Only engines measured on every register appear.

Where we still lose, and why

We lead both means, and both metrics on most registers. Each register we do not lead outright has a named cause.

The two Calfa models that still lead somewhere are both CC BY-NC 4.0: we can measure them, but nobody can ship them in a commercial pipeline. Ours can be: the code and weights are Apache 2.0, and the word list is CC BY-SA 4.0, as the sources it is counted from require.

What we would tell ourselves a month ago

Read the metric before reading the model. We spent a brief planning to train our way out of a character-similarity gap that a one-word flag had largely created, and nearly built a dictionary harvest for a register whose losses turned out to be its index pages’ dashes. Each of these was found by looking at one page’s output beside its transcript, not by looking at another number.