Licensing
This repository holds three kinds of material with different rights, so one licence over everything would misdescribe at least one of them. Copyright in the code and documentation is held jointly by Stephen Masters and Yvette Mankerian.
| Material | Licence |
|---|---|
Code — src/, tests/, tools/, evaluation/ocr/harness.py | MIT |
Documentation — docs/, README.md, and the benchmark results in evaluation/ | CC BY 4.0 |
Ground-truth transcripts — evaluation/ocr/corpus/expected/ | CC BY 4.0 |
Corpus images — evaluation/ocr/corpus/images/ | Public domain — see below |
The table covers this repository only. The Armenian recognition model is not here and is licensed separately — see below.
The corpus images
Every image in evaluation/ocr/corpus/images/ is in the public domain in the
United States. We did not create them and hold no copyright in them, so there
is nothing here for us to license — a licence file covering them would
overclaim. Provenance, source institution and rights status for each item are
recorded in evaluation/ocr/corpus/SOURCES.md.
Public domain status is jurisdictional. These items are public domain in the US; if you are republishing elsewhere, check your own position.
The transcripts
The ground-truth transcripts in evaluation/ocr/corpus/expected/ were generated by
Claude from the corpus images and are released under CC BY 4.0. They are the
reference the whole benchmark is scored against, and the most directly reusable
artefact here — if you use them, attribution is the only condition.
Bear in mind what they are: a strong vision model’s reading of these images, not a human-verified transcription. The method page sets out the limits.
Attribution
For the documentation, the benchmark results or the transcripts:
Tetrak OCR (Stephen Masters and Yvette Mankerian), CC BY 4.0. https://github.com/scattercode/tetrak
The Armenian recognition model
The tetrak_hy Armenian recogniser is not in this repository and is not
covered by the table above. No weights, training crops or model artefacts are
held here; Tetrak consumes the model as an ordinary installed dependency, the
same way it consumes every other engine.
| Artefact | Where it lives | Licence |
|---|---|---|
| Trained weights | tetrak/easyocr-armenian on Hugging Face | Apache 2.0 |
| The library that ships them | tetrak-easyocr-armenian on PyPI | Apache 2.0 |
| The trainer that produces them | tetrak-hy-trainer | Apache 2.0 |
| Training data | tetrak/armenian-ocr-crops on Hugging Face | CC BY-SA 4.0 |
| Word list for lexicon-aided decoding | wordlist.tsv.gz beside the weights on Hugging Face | CC BY-SA 4.0 |
Apache 2.0 rather than MIT because those repositories derive from
Apache-licensed upstreams (EasyOCR, and Clova AI’s
deep-text-recognition-benchmark beneath it).
The weights are Apache 2.0 while their training text is CC BY-SA 3.0 — pages of the Armenian Soviet Encyclopedia from Armenian Wikisource. Whether share-alike reaches through training data into model parameters is unsettled and jurisdiction-dependent; Creative Commons’ own guidance says so, and recommends share-alike on such models as a conservative position rather than as settled law. We took the permissive reading deliberately, and the training sources and their licences are named on the model card. That reasoning is recorded in full in this project’s decision records, not restated here.
The word list is a different case, and takes the conservative reading. It is
counted directly from CC BY-SA 3.0 Wikisource text, with every word form of
the Nayiri Armenian Lexicon
(© Serouj Ourishian, CC BY 4.0) added, so it is licensed CC BY-SA 4.0: a
version both sources permit. Its attribution is recorded in the release’s
provenance.json and in the library’s NOTICE.
Dependencies
The core dependencies are permissively licensed — Pillow, pytesseract
(Apache-2.0), pdf2image (MIT) and python-dotenv (BSD-3-Clause) — so installing
tetrak-ocr on its own carries no copyleft obligation.
The optional backends are a different matter, and one is worth knowing about
before you build on it: marker-pdf and its Surya models are, as of writing,
GPL-3.0 with a commercial-use exception offered by their author. That does not
affect this project’s licence — the extras are optional dependencies, resolved
at install time and never vendored here — but it does affect anyone
redistributing a bundle that includes them. Check the current terms upstream
rather than relying on this note; licence terms change.