Tetrak OCR
GitHub ↗
Reference

Licensing

This repository holds three kinds of material with different rights, so one licence over everything would misdescribe at least one of them. Copyright in the code and documentation is held jointly by Stephen Masters and Yvette Mankerian.

MaterialLicence
Code — src/, tests/, tools/, evaluation/ocr/harness.pyMIT
Documentation — docs/, README.md, and the benchmark results in evaluation/CC BY 4.0
Ground-truth transcripts — evaluation/ocr/corpus/expected/CC BY 4.0
Corpus images — evaluation/ocr/corpus/images/Public domain — see below

The table covers this repository only. The Armenian recognition model is not here and is licensed separately — see below.

The corpus images

Every image in evaluation/ocr/corpus/images/ is in the public domain in the United States. We did not create them and hold no copyright in them, so there is nothing here for us to license — a licence file covering them would overclaim. Provenance, source institution and rights status for each item are recorded in evaluation/ocr/corpus/SOURCES.md.

Public domain status is jurisdictional. These items are public domain in the US; if you are republishing elsewhere, check your own position.

The transcripts

The ground-truth transcripts in evaluation/ocr/corpus/expected/ were generated by Claude from the corpus images and are released under CC BY 4.0. They are the reference the whole benchmark is scored against, and the most directly reusable artefact here — if you use them, attribution is the only condition.

Bear in mind what they are: a strong vision model’s reading of these images, not a human-verified transcription. The method page sets out the limits.

Attribution

For the documentation, the benchmark results or the transcripts:

Tetrak OCR (Stephen Masters and Yvette Mankerian), CC BY 4.0. https://github.com/scattercode/tetrak

The Armenian recognition model

The tetrak_hy Armenian recogniser is not in this repository and is not covered by the table above. No weights, training crops or model artefacts are held here; Tetrak consumes the model as an ordinary installed dependency, the same way it consumes every other engine.

ArtefactWhere it livesLicence
Trained weightstetrak/easyocr-armenian on Hugging FaceApache 2.0
The library that ships themtetrak-easyocr-armenian on PyPIApache 2.0
The trainer that produces themtetrak-hy-trainerApache 2.0
Training datatetrak/armenian-ocr-crops on Hugging FaceCC BY-SA 4.0
Word list for lexicon-aided decodingwordlist.tsv.gz beside the weights on Hugging FaceCC BY-SA 4.0

Apache 2.0 rather than MIT because those repositories derive from Apache-licensed upstreams (EasyOCR, and Clova AI’s deep-text-recognition-benchmark beneath it).

The weights are Apache 2.0 while their training text is CC BY-SA 3.0 — pages of the Armenian Soviet Encyclopedia from Armenian Wikisource. Whether share-alike reaches through training data into model parameters is unsettled and jurisdiction-dependent; Creative Commons’ own guidance says so, and recommends share-alike on such models as a conservative position rather than as settled law. We took the permissive reading deliberately, and the training sources and their licences are named on the model card. That reasoning is recorded in full in this project’s decision records, not restated here.

The word list is a different case, and takes the conservative reading. It is counted directly from CC BY-SA 3.0 Wikisource text, with every word form of the Nayiri Armenian Lexicon (© Serouj Ourishian, CC BY 4.0) added, so it is licensed CC BY-SA 4.0: a version both sources permit. Its attribution is recorded in the release’s provenance.json and in the library’s NOTICE.

Dependencies

The core dependencies are permissively licensed — Pillow, pytesseract (Apache-2.0), pdf2image (MIT) and python-dotenv (BSD-3-Clause) — so installing tetrak-ocr on its own carries no copyleft obligation.

The optional backends are a different matter, and one is worth knowing about before you build on it: marker-pdf and its Surya models are, as of writing, GPL-3.0 with a commercial-use exception offered by their author. That does not affect this project’s licence — the extras are optional dependencies, resolved at install time and never vendored here — but it does affect anyone redistributing a bundle that includes them. Check the current terms upstream rather than relying on this note; licence terms change.