Notes on the state of archive digitisation and OCR, and on what is changing inside Tetrak.
Two threads run through this section, and they are the same thread seen from
either end.
The first is the field: what OCR can and cannot currently do with the material
that actually sits in archives — the linen postcards, the multi-column trade
journals, the hand-lettered posters — and how the people who look after those
collections are getting text out of them today.
The second is this project. Where the routing has been changed and why, what the
quality gate is now catching that it missed before, and how the Armenian EasyOCR
model is coming along. Progress reports, including the ones where the number
went the wrong way.
Nothing here restates the benchmark. The research is where the
measured comparison lives, and it is regenerated from the harness rather than
written; these are the arguments and the working notes around it.
Ten human-proofread encyclopedia pages, ten OCR configurations measured against them, and the first honest numbers from a recogniser we are training ourselves — with the code, data and weights all published in the open