Python API
Resolve a backend by name rather than importing one directly — the registry
reports a missing optional dependency as an actionable error instead of
ModuleNotFoundError.
from pathlib import Path
from tetrak_ocr.registry import get_backend
ocr = get_backend("tesseract-auto")
text = ocr(Path("scan.jpg"))
Registry
tetrak_ocr.registry
Resolve a backend by name.
Every backend exposes the same callable — ocr_image(path) -> str — so the
rest of the package can treat them interchangeably. This module is where a
name becomes one of those callables, and where a missing optional dependency
turns into an error that says which extra to install.
Importing a backend module is always safe: each one records whether its heavy
dependency imported successfully in _IMPORT_OK rather than failing at
import time. That matters because :func:available needs to inspect every
backend, including the ones this machine cannot run.
get_backend(name: 'str') -> 'Callable[..., str]'
Return the ocr_image callable for name.
Raises :class:UnknownBackendError for a name that does not exist, and
:class:MissingBackendError when the backend exists but its extra is not
installed.
is_available(name: 'str') -> 'bool'
True when name can actually run on this machine.
available() -> 'list[str]'
Every backend whose dependencies are installed here.
Used by the CLI and the evaluation harness to skip backends the current environment cannot run, instead of failing the whole run.
supported_extensions(name: 'str') -> 'set[str]'
File extensions name accepts, from the backend’s own declaration.
Accuracy metrics
tetrak_ocr.accuracy
Text similarity metrics for OCR evaluation.
Two metrics are provided:
character_similarity Uses Python’s difflib.SequenceMatcher to compare the two texts character-by-character. Returns a ratio from 0.0 (nothing in common) to 1.0 (identical after normalisation). Good for detecting small transcription errors and OCR noise.
word_recall Computes what fraction of the words in the expected text were also found in the actual OCR output. This is a recall-oriented metric: it rewards capturing all the content, and is tolerant of reordering or extra words introduced by the OCR engine.
Both functions normalise their inputs first (lowercase, collapsed whitespace) so that trivial formatting differences do not affect scores.
normalise(text: str) -> str
Lowercase, equate Armenian punctuation homoglyphs, collapse whitespace.
Two Armenian marks are scored as the ASCII character they print
identically to: the full stop ։ as the colon, and the abbreviation
dot ․ (U+2024) as the full stop. Transcribers type the ASCII form
so often (the medical encyclopedia’s transcripts use the colon
throughout; every register uses . for ․) that holding an engine
to either form would rank engines on which habit their output happens
to share with the transcript, not on what they read. Mapped towards
ASCII so that text with no Armenian in it is unaffected.
character_similarity(actual: str, expected: str) -> float
Return character-level similarity as a ratio 0.0–1.0.
Uses difflib.SequenceMatcher, which finds the longest common subsequences between the two strings.
Args: actual: The text produced by OCR. expected: The reference (ground-truth) text.
Returns: A float in [0.0, 1.0]. 1.0 means the texts are identical after normalisation.
autojunk is off, and must stay off. With difflib’s default, any
string over 200 characters has every character making up more than 1%
of it treated as junk – on a page of text, most of the alphabet and
the space – so the ratio reflects where the few surviving matches
happen to fall rather than how much of the text was read. Brief 013
found it scoring a near-perfect Armenian page at 0.76 instead of 0.98.
word_recall(actual: str, expected: str) -> float
Return the fraction of expected words present in the OCR output.
This measures recall: how much of the expected content did we capture? It is tolerant of the OCR engine producing extra words or changing word order, which is common with complex layouts.
Args: actual: The text produced by OCR. expected: The reference (ground-truth) text.
Returns: A float in [0.0, 1.0]. 1.0 means every expected word was found.
Quality scoring
tetrak_ocr.qa_score
Reference-free OCR quality metrics.
Two metrics are combined into a single score used to rank competing transcripts produced by different OCR backends, and to decide whether a document should be routed to the triage queue.
dictionary_coverage Fraction of word tokens that are recognised English words, using pyspellchecker (pure Python, ~100 KB, no model downloads). Effective at catching non-word OCR noise (“tbe”, “h0use”, “rece1ved”).
perplexity_score GPT-2 language model perplexity via the Hugging Face transformers library. Lower values mean more coherent English prose. The GPT-2 model (~500 MB) is downloaded on first use and cached at ~/.cache/huggingface/. The model and tokeniser are kept as module-level singletons so the cost is paid once per process.
combined_score dict_coverage × (1 / log1p(perplexity)). Higher is better.
LowQualityError Raised by tetrak_ocr.auto_local when the best backend’s combined_score falls below MIN_QUALITY_THRESHOLD. Carries the per-backend scores and raw transcripts so tetrak_ocr.batch can write a useful triage manifest.
These metrics read English, and only English. dictionary_coverage
counts [a-zA-Z] tokens against an English wordlist and GPT-2 is an English
language model, so a perfectly good transcript in another script scores 0.0
and lands below any threshold. Use scores_english_text before applying the
gate to material that may not be English – see tetrak-ocr batch --quality-gate, which refuses rather than triaging a whole run.
LowQualityError(message: str, scores: dict[str, float], transcripts: dict[str, str]) -> None
Raised when all OCR backends produce output below MIN_QUALITY_THRESHOLD.
Attributes: scores: {backend_name: combined_score} for every backend tried. transcripts: {backend_name: text} for every backend tried.
is_available() -> bool
True when every package the scoring path imports is installed.
scores_english_text(text: str) -> bool
Whether these metrics can say anything about text.
Both are English: dictionary_coverage counts [a-zA-Z] tokens against an
English wordlist, and GPT-2 models English prose. An Armenian transcript
therefore scores 0.0 regardless of how well it was read – there is nothing
in it for either metric to look at.
That matters because 0.0 is below every threshold. Gating on it would send
a whole run to triage while reporting each file as poor quality, which is
the opposite of the truth and gives no clue where to look. The
easyocr-hy backend made this reachable; before it, everything the
pipeline read was English.
A transcript with no letters at all returns True: empty and letterless output is poor, and scoring says so correctly. Only a transcript in another script is the case this guards.
dictionary_coverage(text: str) -> float
Return the fraction of word tokens that are recognised English words.
Tokens shorter than 2 characters are excluded (they are too short to be meaningful spell-check targets and common as OCR noise).
Args: text: Raw text string.
Returns: A float in [0.0, 1.0]. Returns 0.0 for empty or token-free input.
perplexity_score(text: str) -> float
Return the GPT-2 perplexity of the text.
Lower values indicate more coherent, English-like prose. Typical ranges:
- Clean prose: 20–80
- Moderate errors: 100–500
- Heavy OCR noise: 500+
Text is truncated to 1024 tokens (GPT-2 context limit). Very short texts (fewer than 5 tokens) return a fixed high value (1000.0) since perplexity is unreliable on such short sequences.
Args: text: Raw text string.
Returns: A positive float. Lower is better.
combined_score(text: str) -> float
Return a single quality score combining dictionary coverage and perplexity.
Score = dict_coverage × (1 / log1p(perplexity))
Higher is better. log1p tames the unbounded perplexity scale. Returns 0.0 if text is empty.
Args: text: Raw text string.
Returns: A non-negative float. Higher means better quality.
Image and frame handling
tetrak_ocr.imaging
Frame handling for multi-page raster images.
TIFF is the one raster format in SUPPORTED_EXTENSIONS that can hold more than one page, and archival TIFF frequently does — a scanned pamphlet or register often arrives as a single file with one frame per leaf.
Pillow opens such a file at frame 0 and says nothing about the rest. Every backend here loads images through Pillow, directly or indirectly, so “TIFF support” meant reading page one and silently discarding the others. That is the same failure the Claude backend already guards against with its max_tokens check: a truncated transcript is worse than a failed one, because it looks like a result and quietly corrupts everything downstream of it.
So the rule in this package is that a multi-frame image is either read in full
or refused by name. Tesseract reads it in full, page by page, exactly as it
already does for PDFs. The backends that cannot yet do so raise
:class:MultiPageNotSupportedError, which names the file, the page count and a
backend that will read it — rather than returning page one as though that were
the document.
pdf_renderer()
Return pdf2image.convert_from_path, or explain why it is missing.
Every PDF path in the package goes through here so the advice is given
once and is correct. pdf2image is a core dependency, not an extra, so
an ImportError means an incomplete install – the three call sites used to
say pip install -r requirements.txt, naming a file this project has not
had since it became a package, so the one instruction offered could not
work.
frame_count(path: 'Path') -> 'int'
Number of pages in a raster image; 1 for ordinary single-page files.
Pages, not IFDs. Not every extra frame in a TIFF is another leaf of the
document – see :func:_is_page.
Non-raster inputs (PDFs) and anything Pillow cannot open report 1: this is a question about TIFF paging, and callers handle those cases by other routes. It never raises, so it is safe to call as a guard.
page_count(path: 'Path') -> 'int'
Pages in a document of any supported kind, without rendering it.
:func:frame_count answers this for rasters but reports 1 for a PDF,
because paging a PDF is somebody else’s route. This closes that gap for
callers who want the number rather than the pages – telemetry, mainly,
where “how long did it take” is meaningless next to a page count.
PDFs go through poppler’s pdfinfo, which reads the trailer rather than
rasterising; :func:iter_pages would render every page to count them.
Never raises: a page count is not worth an exception mid-batch.
iter_frames(path: 'Path') -> 'Iterator[Image.Image]'
Yield every page of path as an independent image.
Reduced-resolution renditions and transparency masks are skipped; see
:func:_is_page.
The frame’s mode is preserved, not normalised. A bilevel TIFF yields
mode “1”, a palette one yields “P”. That is deliberate: converting to RGB
first was measured and changes no OCR output, so it would be a conversion
that costs memory and buys nothing. preprocess greyscales whatever it is
given. See TestFrameModeIsNotAProblem for the measurement.
Each frame is copied before being yielded. Pillow’s frames are views onto
one open file that seek mutates in place, so a caller that collected them
without copying would end up holding several references to the last frame.
reject_multi_page(path: 'Path', backend: 'str', *, use_instead: 'str' = 'tesseract') -> 'None'
Raise if path holds more than one frame.
For backends that read only frame 0. Called before any work is done, so the caller learns the file cannot be read properly rather than receiving a plausible-looking transcript of its first page.
iter_pages(path: 'Path') -> 'Iterator[Image.Image]'
Yield every page of a document, whether it is a PDF or a raster.
:func:iter_frames handles multi-frame rasters and PDFs are handled by
pdf2image – two routes to the same idea, which the Tesseract backend
has always kept apart because it only ever needed the text.
Producing a searchable PDF needs both halves to agree: the OCR pass reads the pages to find where the words are, and the writer reads them again to draw them. If those two disagreed about what page three is, the text layer would land on the wrong image. Going through one iterator is what stops that being possible.
streamed_pages(path: 'Path') -> 'Iterator[Iterator[Image.Image]]'
Yield page images one at a time, without holding the document in memory.
:func:iter_pages looks like a generator but is not a streaming one for
PDFs: convert_from_path renders every page into a list before the first
is yielded, so a 21-page document costs 881 MB of peak RSS whether the
caller wants one page or all of them.
Rendering into a temporary directory instead hands back images backed by
files, which Pillow loads on demand. Same single poppler invocation, same
wall time (13.8s against 13.7s measured), 58 MB peak instead of 881.
Rendering a page at a time with first_page/last_page also bounds
memory but costs an invocation per page (90 MB, 14.9s), so it loses on both
counts.
The images are only valid inside the block – their backing files are deleted on exit. Close each one when done with it, or the loaded pixel data accumulates and the saving is undone.
Multi-frame rasters need none of this: :func:iter_frames already yields
one frame at a time.
page_documents(path: 'Path') -> 'Iterator[list[Path]]'
Split path into one single-page document per page, in a temp directory.
Yields the page paths in order; they are deleted on exit, so a caller that needs the transcripts must read them inside the block.
PDFs are split as PDFs, never rendered to images. Marker reads a PDF’s existing text layer through pdftext instead of OCR-ing it, and rendering a page to PNG throws that layer away: measured at 1.8x slower and losing text outright, with one page dropping from 167 characters to zero. See design research note 001 (Marker throughput and multi-page PDFs).
Multi-frame rasters have no text layer to preserve, so their frames are
written as single-frame TIFFs – lossless, and preserving the frame mode
that :func:iter_frames deliberately does not normalise.
A single-page document yields itself unchanged rather than a copy: there is nothing to split, and the copy would only be a slower way to say so.
Auto-local routing
tetrak_ocr.auto_local
auto-local OCR backend: fan out over the local engines and keep the best.
Runs every eligible local backend, scores each transcript with qa_score.combined_score() (dictionary coverage x inverse log-perplexity), and returns the highest-ranked one. It exists because no single local engine wins across document types, so choosing per file beats choosing per batch.
The winning transcript passes a quality gate: if it falls below
qa_score.MIN_QUALITY_THRESHOLD, LowQualityError is raised so the caller can
route the document to the triage queue rather than silently writing a bad
result. The same gate is available for any single backend via
tetrak-ocr batch --quality-gate; this one additionally carries the
per-backend score table, which is what makes the triage manifest worth
reading.
Backend eligibility:
GPU available + Marker installed: Images: Marker, Vision, EasyOCR, Paddle, Tesseract-auto (each if installed) PDFs: Marker, Tesseract-auto
No GPU: Images: Vision, EasyOCR, Paddle, Tesseract-auto (each if installed) PDFs: Tesseract-auto
With --with-paddle-vl, PaddleOCR-VL joins both lists, images and PDFs
alike. It is opt-in rather than automatic because it runs in minutes where
the rest of the pool runs in seconds – see _build_candidates().
Vision, EasyOCR and Paddle cannot read PDFs, so all three are excluded for those. Vision is additionally macOS-only; on any other platform its extra cannot install and it simply never appears as a candidate.
Fan-out runs its backends sequentially, by design. Running them in parallel looks like free speed and is not: these engines are individually heavy on CPU and RAM, and several hold module-level model singletons. Starting two or three at once on one machine risks memory pressure and contention that would make runtimes less predictable, not more – which matters most on exactly the large batch jobs where the time would otherwise be worth saving.
There was briefly a second strategy here, auto-local-fast, which ran only the
top-ranked eligible backend. The benchmark retired it: in every installed
configuration it was at best equal to running tesseract-auto directly, and on
a GPU machine it resolved to Marker and cost twenty-four times as much for less
accuracy. If a single cheap engine is what you want, name it – and add
--quality-gate if you want the triage queue with it.
Public API: has_gpu() -> bool — True if CUDA, ROCm, or Apple MPS is detected ocr_image(path) -> str — OCR a file; raises LowQualityError if poor SUPPORTED_EXTENSIONS — set of supported file extensions
Model weights are downloaded on first use and cached locally. Singleton patterns (same as EasyOCR, PaddleOCR) ensure each model loads only once.
has_gpu() -> bool
Return True if a GPU is available for neural-network acceleration.
Detects NVIDIA CUDA, AMD ROCm, and Apple Metal Performance Shaders (MPS). Returns False if torch is not importable or no supported GPU is found.
ocr_image(path: 'Path | str', with_paddle_vl: bool = False) -> str
OCR a file by running every eligible local backend and keeping the best.
Args:
path: Path to an image or PDF file.
with_paddle_vl: Admit the PaddleOCR-VL backend to the pool. Off by
default because it costs minutes per page against the rest of
the pool’s seconds; see :func:_build_candidates. The CLI
exposes it as --with-paddle-vl. An optional keyword rather
than a change to the ocr_image(path) -> str contract, the
same way the Tesseract backend takes contrast and psm.
Fan-out exists because no single local backend wins across document types,
so choosing per file beats choosing per batch – see
tetrak-ocr evaluate --all for the measurement behind that.
The winning transcript passes a quality gate, so a file that no backend reads acceptably reaches the triage queue instead of being written out.
Returns: Extracted text as a string.
Raises: LowQualityError: If the winning combined score falls below MIN_QUALITY_THRESHOLD. The exception carries .scores and .transcripts dicts for triage queue reporting.
Output formats
tetrak_ocr.outputs
Output formats for a finished transcript.
One transcript can be written several ways at once — Markdown for a human, plain text for a downstream tool, a searchable PDF for a catalogue — so the pipeline asks for a set of formats rather than a single one.
Every writer exposes the same callable::
write(source, transcript, destination, **options) -> None
source is the file that was transcribed, which the text formats ignore and
the PDF writer needs (it embeds the scan). options carries format-specific
settings; a writer must tolerate options meant for a different format, because
the caller passes one dict to all of them.
Adding a format means adding one entry to _FORMATS and nothing else: the
CLI takes its choices, its help text and its validation from this table, the
same way backends are resolved through registry.py.
canonical(name: 'str') -> 'str'
Resolve a format name or alias, raising for anything unrecognised.
parse(values: 'Iterable[str] | None') -> 'list[str]'
Turn CLI --format values into an ordered list of canonical names.
Accepts both spellings people try — repeated flags and one comma-separated value — because guessing wrong about which a tool supports is a papercut::
--format markdown --format pdf
--format markdown,pdf
Order follows first mention, and duplicates collapse, so asking for
markdown,markdown,pdf is not an error and does not write twice.
Returns the default when nothing was asked for.
extension(name: 'str') -> 'str'
The file extension a format writes, including the leading dot.
describe() -> 'str'
Format list for --help, one per line.
write(name: 'str', source: 'Path', transcript: 'str', destination_dir: 'Path', **options) -> 'Path'
Write transcript in format name, returning the path written.
The destination filename is the source’s stem plus the format’s extension, which is what keeps a document’s outputs sitting together under one name.
Except when that would collide with the source itself. A PDF
transcribed to a searchable PDF wants the same name as the original, and
the batch pipeline moves the original into the same directory – so one
silently overwrote the other, with the run reporting success either way.
Such an output is written as <stem>.searchable<ext> instead. Only a
source whose own extension is also an output format can hit this, which
today means PDF alone.
Telemetry
tetrak_ocr.telemetry
Append-only run log: what was processed, by which backend, and how long it took.
A long batch is otherwise silent about its own cost. auto-local prints its
score table only once every backend has finished, so a file that takes forty
minutes looks identical to one that has hung – and when it does finish, the
table says which backend won but not which one spent the time.
This module records that as it happens. Every event is one JSON object on one
line, flushed immediately, so tail -f on the log is a live progress view and
the finished file is a dataset:
{"ts": "...", "run": "20260822T115103Z-7f3a", "event": "backend_finish",
"file": "camera-1919.pdf", "backend": "marker", "seconds": 2246.8, ...}
JSON Lines rather than CSV because the fields differ per event and will grow; rather than a database because appending a line is atomic enough for a sequential pipeline and needs nothing installed to read it back.
The sink is process-global. That is deliberate: backends expose a fixed
ocr_image(path) -> str signature with nowhere to thread a logger through,
and the fan-out is documented as sequential, so there is no concurrent writer
to coordinate. When no run has been started every call is a cheap no-op, which
is what keeps the library usable as a library.
TelemetryLog(path: 'Path', run_id: 'str') -> 'None'
One run’s worth of events, appended to a JSON Lines file.
start_run(path: 'Path | str | None' = None, **fields) -> 'TelemetryLog | None'
Begin recording, returning the log (or None when disabled).
finish_run(**fields) -> 'None'
Close out the current run and detach the sink.
record(event: 'str', **fields) -> 'None'
Emit an event if a run is active, otherwise do nothing.
active() -> 'bool'
No docstring. This function is public and undocumented.
log_path() -> 'Path | None'
No docstring. This function is public and undocumented.
describe_file(path: 'Path') -> 'dict'
The file facts worth having on every record about it.
Page count matters more than byte size here and is the thing people are surprised by: a 3 MB PDF that is ten pages is ten documents of work, and reading a per-file duration without it invites the wrong conclusion. It is read cheaply and never allowed to fail – a page count is not worth an exception in the middle of a batch.
timed(event: 'str', **fields)
Time a block and emit event with its duration, success or not.
Emits on the failure path too, because a backend that raised after twenty minutes is exactly the thing worth having in the log.
Batch pipeline
tetrak_ocr.batch
Batch OCR processor.
Scans the workspace/scans/ directory for supported image and PDF files, runs
OCR on each one, writes the extracted text to a Markdown file in
workspace/processed/, and moves the original file in alongside it.
After processing, each file pair looks like: workspace/processed/ my-postcard.jpg ← original moved here my-postcard.md ← extracted text
Usage (run from the repository root): tetrak-ocr batch tetrak-ocr batch –backend auto-local tetrak-ocr batch –backend claude tetrak-ocr batch –backend tesseract –contrast 3.0
build_ocr_fn(backend: str, contrast: float | None = None, psm: int | None = None, auto: bool = False, with_paddle_vl: bool = False) -> collections.abc.Callable[[pathlib.Path], str]
Return the OCR callable for the chosen backend.
Backend resolution lives in :mod:tetrak_ocr.registry — this used to
duplicate it as an if/elif chain of imports, which meant a new backend had
to be registered in two places and a missing optional dependency surfaced
as a bare ImportError.
Only Tesseract takes tuning arguments and only auto-local takes
--with-paddle-vl; the rest ignore them, and passing one is reported
rather than silently dropped.
Both Tesseract names take the flags. tesseract-auto is the same backend
with auto-configuration switched on by the registry, so it honours
--contrast and --psm for whichever of the two you name; it used to drop
them, because the wiring below tested == "tesseract" where the warning
above tests startswith, leaving that one name in neither branch.
A value left as None is what lets auto-configuration fill it.
analyse_image() supplies exactly the parameters the caller did not, so
substituting a default here silently turns auto off – which is what made
--auto inert on the tesseract backend: contrast and psm were always
filled before the backend could look at the image.
supported_extensions(backend: str) -> set[str]
File extensions backend accepts.
auto-local routes PDFs to Tesseract, so it reports Tesseract’s full set — the registry reads this from each backend’s own SUPPORTED_EXTENSIONS.
transcribe_per_page(image_path: pathlib.Path, ocr_fn: collections.abc.Callable[[pathlib.Path], str]) -> list[str]
Transcribe image_path one page at a time, returning a text per page.
Needed only for a multi-page searchable PDF, which has to lay each page’s words onto the page they came from. It costs roughly twice a whole-document pass – measured, and not caused by anything we can avoid: reusing one Marker converter across the pages made no difference. See design research note 001 (Marker throughput and multi-page PDFs).
That is why the caller takes this route only when the pdf format is
actually asked for on a multi-page document, rather than always.
process_file(image_path: pathlib.Path, ocr_fn: collections.abc.Callable[[pathlib.Path], str], triaged: list[str] | None = None, quality_gate: bool = False, backend: str | None = None, formats: list[str] | None = None, pdf_full_size: bool = False, pdf_font: str | None = None) -> None
OCR a single file, write Markdown output, and move it to workspace/processed/.
A transcript that scores below MIN_QUALITY_THRESHOLD goes to the triage queue instead of workspace/processed/, so a bad read is triaged rather than silently written out. There are two routes to that, because there are two things that know a transcript is poor:
auto-localscores every backend it ran in order to pick a winner, so it raises LowQualityError itself and hands over the whole score table. That table is what makes the triage manifest worth reading.- Any other backend produces one transcript and no opinion about it. When quality_gate is set, this function scores that transcript and applies the same threshold.
The gate is opt-in for single backends rather than always on: it changes
which files reach workspace/processed/, and turning it on silently would start
diverting output for anyone already running --backend tesseract.
Args: image_path: Path to the file inside workspace/scans/. ocr_fn: The OCR callable to use. triaged: Optional list to append filenames routed to triage. quality_gate: Score this backend’s transcript and gate on it. backend: Backend name, used to label the triage manifest. formats: Output formats to write, defaulting to Markdown alone. Every requested format is written for the same transcript – the OCR runs once regardless of how many are asked for.
main(args=None) -> int
Parse arguments and process all supported files found in workspace/scans/.
Returns a process exit code – 0 on success, 1 if any file failed or
workspace/scans/ is missing – so that callers can propagate a partial failure
instead of reporting success. Every failure this function decides on is
returned rather than raised, which is what lets cli.main honour its own
-> int contract.
The one exception is not ours: argparse raises SystemExit on an
unparseable argument list before this function regains control. Callers
driving it programmatically with untrusted arguments should expect that.
Errors
tetrak_ocr.errors
Exceptions raised by the package.
These exist mainly so that a missing optional backend fails as a library
should — with a catchable exception naming the extra to install — rather than
by printing to stderr and calling sys.exit(1), which is what the original
scripts did. That was reasonable when each file was run directly; it is not,
when a caller imports one backend and wants to fall back to another.
OcrPipelineError(…)
Base class for every error this package raises.
MissingBackendError(backend: 'str', extra: 'str', packages: 'str') -> 'None'
A backend was requested but its optional dependency is not installed.
The message names the extra to install, because “No module named ‘paddleocr’” tells a user what broke but not what to do about it.
UnknownBackendError(name: 'str', known: 'list[str]') -> 'None'
A backend name was requested that does not exist.
UnknownFormatError(name: 'str', valid: 'list[str]') -> 'None'
An output format was requested that does not exist.
Carries the valid names: the useful thing to tell someone who typed
--format pdff is what they could have typed instead.
MultiPageNotSupportedError(path: 'Path', backend: 'str', pages: 'int', use_instead: 'str') -> 'None'
A multi-frame image was given to a backend that reads only frame 0.
Raised instead of transcribing page one and returning it as though it were the whole document. Archival TIFF is often multi-page, and a partial transcript that looks complete is the failure mode this package treats as worst: it passes quality scoring, enters the archive, and misleads every reader after that.
MultiPagePdfError(path: 'Path', pages: 'int') -> 'None'
A searchable PDF was asked for on a document with more than one page.
The text layer carries the transcript the pipeline chose, and a backend returns one transcript for a whole document rather than one per page. There is no honest way to divide it: the pages would be guesses, and a reader searching a 21-page programme would be sent to the wrong leaf while the file looked entirely correct.
So this refuses, for the same reason
:class:MultiPageNotSupportedError does. Per-page transcripts arrive with
the positioned text layer – see
brief 008 (searchable PDF).