Tetrak OCR
GitHub ↗
Reference

Python API

Resolve a backend by name rather than importing one directly — the registry reports a missing optional dependency as an actionable error instead of ModuleNotFoundError.

from pathlib import Path
from tetrak_ocr.registry import get_backend

ocr = get_backend("tesseract-auto")
text = ocr(Path("scan.jpg"))

Registry

tetrak_ocr.registry

Resolve a backend by name.

Every backend exposes the same callable — ocr_image(path) -> str — so the rest of the package can treat them interchangeably. This module is where a name becomes one of those callables, and where a missing optional dependency turns into an error that says which extra to install.

Importing a backend module is always safe: each one records whether its heavy dependency imported successfully in _IMPORT_OK rather than failing at import time. That matters because :func:available needs to inspect every backend, including the ones this machine cannot run.

get_backend(name: 'str') -> 'Callable[..., str]'

Return the ocr_image callable for name.

Raises :class:UnknownBackendError for a name that does not exist, and :class:MissingBackendError when the backend exists but its extra is not installed.

is_available(name: 'str') -> 'bool'

True when name can actually run on this machine.

available() -> 'list[str]'

Every backend whose dependencies are installed here.

Used by the CLI and the evaluation harness to skip backends the current environment cannot run, instead of failing the whole run.

supported_extensions(name: 'str') -> 'set[str]'

File extensions name accepts, from the backend’s own declaration.

Accuracy metrics

tetrak_ocr.accuracy

Text similarity metrics for OCR evaluation.

Two metrics are provided:

character_similarity Uses Python’s difflib.SequenceMatcher to compare the two texts character-by-character. Returns a ratio from 0.0 (nothing in common) to 1.0 (identical after normalisation). Good for detecting small transcription errors and OCR noise.

word_recall Computes what fraction of the words in the expected text were also found in the actual OCR output. This is a recall-oriented metric: it rewards capturing all the content, and is tolerant of reordering or extra words introduced by the OCR engine.

Both functions normalise their inputs first (lowercase, collapsed whitespace) so that trivial formatting differences do not affect scores.

normalise(text: str) -> str

Lowercase, equate Armenian punctuation homoglyphs, collapse whitespace.

Two Armenian marks are scored as the ASCII character they print identically to: the full stop ։ as the colon, and the abbreviation dot ․ (U+2024) as the full stop. Transcribers type the ASCII form so often (the medical encyclopedia’s transcripts use the colon throughout; every register uses . for ․) that holding an engine to either form would rank engines on which habit their output happens to share with the transcript, not on what they read. Mapped towards ASCII so that text with no Armenian in it is unaffected.

character_similarity(actual: str, expected: str) -> float

Return character-level similarity as a ratio 0.0–1.0.

Uses difflib.SequenceMatcher, which finds the longest common subsequences between the two strings.

Args: actual: The text produced by OCR. expected: The reference (ground-truth) text.

Returns: A float in [0.0, 1.0]. 1.0 means the texts are identical after normalisation.

autojunk is off, and must stay off. With difflib’s default, any string over 200 characters has every character making up more than 1% of it treated as junk – on a page of text, most of the alphabet and the space – so the ratio reflects where the few surviving matches happen to fall rather than how much of the text was read. Brief 013 found it scoring a near-perfect Armenian page at 0.76 instead of 0.98.

word_recall(actual: str, expected: str) -> float

Return the fraction of expected words present in the OCR output.

This measures recall: how much of the expected content did we capture? It is tolerant of the OCR engine producing extra words or changing word order, which is common with complex layouts.

Args: actual: The text produced by OCR. expected: The reference (ground-truth) text.

Returns: A float in [0.0, 1.0]. 1.0 means every expected word was found.

Quality scoring

tetrak_ocr.qa_score

Reference-free OCR quality metrics.

Two metrics are combined into a single score used to rank competing transcripts produced by different OCR backends, and to decide whether a document should be routed to the triage queue.

dictionary_coverage Fraction of word tokens that are recognised English words, using pyspellchecker (pure Python, ~100 KB, no model downloads). Effective at catching non-word OCR noise (“tbe”, “h0use”, “rece1ved”).

perplexity_score GPT-2 language model perplexity via the Hugging Face transformers library. Lower values mean more coherent English prose. The GPT-2 model (~500 MB) is downloaded on first use and cached at ~/.cache/huggingface/. The model and tokeniser are kept as module-level singletons so the cost is paid once per process.

combined_score dict_coverage × (1 / log1p(perplexity)). Higher is better.

LowQualityError Raised by tetrak_ocr.auto_local when the best backend’s combined_score falls below MIN_QUALITY_THRESHOLD. Carries the per-backend scores and raw transcripts so tetrak_ocr.batch can write a useful triage manifest.

These metrics read English, and only English. dictionary_coverage counts [a-zA-Z] tokens against an English wordlist and GPT-2 is an English language model, so a perfectly good transcript in another script scores 0.0 and lands below any threshold. Use scores_english_text before applying the gate to material that may not be English – see tetrak-ocr batch --quality-gate, which refuses rather than triaging a whole run.

LowQualityError(message: str, scores: dict[str, float], transcripts: dict[str, str]) -> None

Raised when all OCR backends produce output below MIN_QUALITY_THRESHOLD.

Attributes: scores: {backend_name: combined_score} for every backend tried. transcripts: {backend_name: text} for every backend tried.

is_available() -> bool

True when every package the scoring path imports is installed.

scores_english_text(text: str) -> bool

Whether these metrics can say anything about text.

Both are English: dictionary_coverage counts [a-zA-Z] tokens against an English wordlist, and GPT-2 models English prose. An Armenian transcript therefore scores 0.0 regardless of how well it was read – there is nothing in it for either metric to look at.

That matters because 0.0 is below every threshold. Gating on it would send a whole run to triage while reporting each file as poor quality, which is the opposite of the truth and gives no clue where to look. The easyocr-hy backend made this reachable; before it, everything the pipeline read was English.

A transcript with no letters at all returns True: empty and letterless output is poor, and scoring says so correctly. Only a transcript in another script is the case this guards.

dictionary_coverage(text: str) -> float

Return the fraction of word tokens that are recognised English words.

Tokens shorter than 2 characters are excluded (they are too short to be meaningful spell-check targets and common as OCR noise).

Args: text: Raw text string.

Returns: A float in [0.0, 1.0]. Returns 0.0 for empty or token-free input.

perplexity_score(text: str) -> float

Return the GPT-2 perplexity of the text.

Lower values indicate more coherent, English-like prose. Typical ranges:

Text is truncated to 1024 tokens (GPT-2 context limit). Very short texts (fewer than 5 tokens) return a fixed high value (1000.0) since perplexity is unreliable on such short sequences.

Args: text: Raw text string.

Returns: A positive float. Lower is better.

combined_score(text: str) -> float

Return a single quality score combining dictionary coverage and perplexity.

Score = dict_coverage × (1 / log1p(perplexity))

Higher is better. log1p tames the unbounded perplexity scale. Returns 0.0 if text is empty.

Args: text: Raw text string.

Returns: A non-negative float. Higher means better quality.

Image and frame handling

tetrak_ocr.imaging

Frame handling for multi-page raster images.

TIFF is the one raster format in SUPPORTED_EXTENSIONS that can hold more than one page, and archival TIFF frequently does — a scanned pamphlet or register often arrives as a single file with one frame per leaf.

Pillow opens such a file at frame 0 and says nothing about the rest. Every backend here loads images through Pillow, directly or indirectly, so “TIFF support” meant reading page one and silently discarding the others. That is the same failure the Claude backend already guards against with its max_tokens check: a truncated transcript is worse than a failed one, because it looks like a result and quietly corrupts everything downstream of it.

So the rule in this package is that a multi-frame image is either read in full or refused by name. Tesseract reads it in full, page by page, exactly as it already does for PDFs. The backends that cannot yet do so raise :class:MultiPageNotSupportedError, which names the file, the page count and a backend that will read it — rather than returning page one as though that were the document.

pdf_renderer()

Return pdf2image.convert_from_path, or explain why it is missing.

Every PDF path in the package goes through here so the advice is given once and is correct. pdf2image is a core dependency, not an extra, so an ImportError means an incomplete install – the three call sites used to say pip install -r requirements.txt, naming a file this project has not had since it became a package, so the one instruction offered could not work.

frame_count(path: 'Path') -> 'int'

Number of pages in a raster image; 1 for ordinary single-page files.

Pages, not IFDs. Not every extra frame in a TIFF is another leaf of the document – see :func:_is_page.

Non-raster inputs (PDFs) and anything Pillow cannot open report 1: this is a question about TIFF paging, and callers handle those cases by other routes. It never raises, so it is safe to call as a guard.

page_count(path: 'Path') -> 'int'

Pages in a document of any supported kind, without rendering it.

:func:frame_count answers this for rasters but reports 1 for a PDF, because paging a PDF is somebody else’s route. This closes that gap for callers who want the number rather than the pages – telemetry, mainly, where “how long did it take” is meaningless next to a page count.

PDFs go through poppler’s pdfinfo, which reads the trailer rather than rasterising; :func:iter_pages would render every page to count them. Never raises: a page count is not worth an exception mid-batch.

iter_frames(path: 'Path') -> 'Iterator[Image.Image]'

Yield every page of path as an independent image.

Reduced-resolution renditions and transparency masks are skipped; see :func:_is_page.

The frame’s mode is preserved, not normalised. A bilevel TIFF yields mode “1”, a palette one yields “P”. That is deliberate: converting to RGB first was measured and changes no OCR output, so it would be a conversion that costs memory and buys nothing. preprocess greyscales whatever it is given. See TestFrameModeIsNotAProblem for the measurement.

Each frame is copied before being yielded. Pillow’s frames are views onto one open file that seek mutates in place, so a caller that collected them without copying would end up holding several references to the last frame.

reject_multi_page(path: 'Path', backend: 'str', *, use_instead: 'str' = 'tesseract') -> 'None'

Raise if path holds more than one frame.

For backends that read only frame 0. Called before any work is done, so the caller learns the file cannot be read properly rather than receiving a plausible-looking transcript of its first page.

iter_pages(path: 'Path') -> 'Iterator[Image.Image]'

Yield every page of a document, whether it is a PDF or a raster.

:func:iter_frames handles multi-frame rasters and PDFs are handled by pdf2image – two routes to the same idea, which the Tesseract backend has always kept apart because it only ever needed the text.

Producing a searchable PDF needs both halves to agree: the OCR pass reads the pages to find where the words are, and the writer reads them again to draw them. If those two disagreed about what page three is, the text layer would land on the wrong image. Going through one iterator is what stops that being possible.

streamed_pages(path: 'Path') -> 'Iterator[Iterator[Image.Image]]'

Yield page images one at a time, without holding the document in memory.

:func:iter_pages looks like a generator but is not a streaming one for PDFs: convert_from_path renders every page into a list before the first is yielded, so a 21-page document costs 881 MB of peak RSS whether the caller wants one page or all of them.

Rendering into a temporary directory instead hands back images backed by files, which Pillow loads on demand. Same single poppler invocation, same wall time (13.8s against 13.7s measured), 58 MB peak instead of 881. Rendering a page at a time with first_page/last_page also bounds memory but costs an invocation per page (90 MB, 14.9s), so it loses on both counts.

The images are only valid inside the block – their backing files are deleted on exit. Close each one when done with it, or the loaded pixel data accumulates and the saving is undone.

Multi-frame rasters need none of this: :func:iter_frames already yields one frame at a time.

page_documents(path: 'Path') -> 'Iterator[list[Path]]'

Split path into one single-page document per page, in a temp directory.

Yields the page paths in order; they are deleted on exit, so a caller that needs the transcripts must read them inside the block.

PDFs are split as PDFs, never rendered to images. Marker reads a PDF’s existing text layer through pdftext instead of OCR-ing it, and rendering a page to PNG throws that layer away: measured at 1.8x slower and losing text outright, with one page dropping from 167 characters to zero. See design research note 001 (Marker throughput and multi-page PDFs).

Multi-frame rasters have no text layer to preserve, so their frames are written as single-frame TIFFs – lossless, and preserving the frame mode that :func:iter_frames deliberately does not normalise.

A single-page document yields itself unchanged rather than a copy: there is nothing to split, and the copy would only be a slower way to say so.

Auto-local routing

tetrak_ocr.auto_local

auto-local OCR backend: fan out over the local engines and keep the best.

Runs every eligible local backend, scores each transcript with qa_score.combined_score() (dictionary coverage x inverse log-perplexity), and returns the highest-ranked one. It exists because no single local engine wins across document types, so choosing per file beats choosing per batch.

The winning transcript passes a quality gate: if it falls below qa_score.MIN_QUALITY_THRESHOLD, LowQualityError is raised so the caller can route the document to the triage queue rather than silently writing a bad result. The same gate is available for any single backend via tetrak-ocr batch --quality-gate; this one additionally carries the per-backend score table, which is what makes the triage manifest worth reading.

Backend eligibility:

GPU available + Marker installed: Images: Marker, Vision, EasyOCR, Paddle, Tesseract-auto (each if installed) PDFs: Marker, Tesseract-auto

No GPU: Images: Vision, EasyOCR, Paddle, Tesseract-auto (each if installed) PDFs: Tesseract-auto

With --with-paddle-vl, PaddleOCR-VL joins both lists, images and PDFs alike. It is opt-in rather than automatic because it runs in minutes where the rest of the pool runs in seconds – see _build_candidates().

Vision, EasyOCR and Paddle cannot read PDFs, so all three are excluded for those. Vision is additionally macOS-only; on any other platform its extra cannot install and it simply never appears as a candidate.

Fan-out runs its backends sequentially, by design. Running them in parallel looks like free speed and is not: these engines are individually heavy on CPU and RAM, and several hold module-level model singletons. Starting two or three at once on one machine risks memory pressure and contention that would make runtimes less predictable, not more – which matters most on exactly the large batch jobs where the time would otherwise be worth saving.

There was briefly a second strategy here, auto-local-fast, which ran only the top-ranked eligible backend. The benchmark retired it: in every installed configuration it was at best equal to running tesseract-auto directly, and on a GPU machine it resolved to Marker and cost twenty-four times as much for less accuracy. If a single cheap engine is what you want, name it – and add --quality-gate if you want the triage queue with it.

Public API: has_gpu() -> bool — True if CUDA, ROCm, or Apple MPS is detected ocr_image(path) -> str — OCR a file; raises LowQualityError if poor SUPPORTED_EXTENSIONS — set of supported file extensions

Model weights are downloaded on first use and cached locally. Singleton patterns (same as EasyOCR, PaddleOCR) ensure each model loads only once.

has_gpu() -> bool

Return True if a GPU is available for neural-network acceleration.

Detects NVIDIA CUDA, AMD ROCm, and Apple Metal Performance Shaders (MPS). Returns False if torch is not importable or no supported GPU is found.

ocr_image(path: 'Path | str', with_paddle_vl: bool = False) -> str

OCR a file by running every eligible local backend and keeping the best.

Args: path: Path to an image or PDF file. with_paddle_vl: Admit the PaddleOCR-VL backend to the pool. Off by default because it costs minutes per page against the rest of the pool’s seconds; see :func:_build_candidates. The CLI exposes it as --with-paddle-vl. An optional keyword rather than a change to the ocr_image(path) -> str contract, the same way the Tesseract backend takes contrast and psm.

Fan-out exists because no single local backend wins across document types, so choosing per file beats choosing per batch – see tetrak-ocr evaluate --all for the measurement behind that.

The winning transcript passes a quality gate, so a file that no backend reads acceptably reaches the triage queue instead of being written out.

Returns: Extracted text as a string.

Raises: LowQualityError: If the winning combined score falls below MIN_QUALITY_THRESHOLD. The exception carries .scores and .transcripts dicts for triage queue reporting.

Output formats

tetrak_ocr.outputs

Output formats for a finished transcript.

One transcript can be written several ways at once — Markdown for a human, plain text for a downstream tool, a searchable PDF for a catalogue — so the pipeline asks for a set of formats rather than a single one.

Every writer exposes the same callable::

write(source, transcript, destination, **options) -> None

source is the file that was transcribed, which the text formats ignore and the PDF writer needs (it embeds the scan). options carries format-specific settings; a writer must tolerate options meant for a different format, because the caller passes one dict to all of them.

Adding a format means adding one entry to _FORMATS and nothing else: the CLI takes its choices, its help text and its validation from this table, the same way backends are resolved through registry.py.

canonical(name: 'str') -> 'str'

Resolve a format name or alias, raising for anything unrecognised.

parse(values: 'Iterable[str] | None') -> 'list[str]'

Turn CLI --format values into an ordered list of canonical names.

Accepts both spellings people try — repeated flags and one comma-separated value — because guessing wrong about which a tool supports is a papercut::

--format markdown --format pdf
--format markdown,pdf

Order follows first mention, and duplicates collapse, so asking for markdown,markdown,pdf is not an error and does not write twice. Returns the default when nothing was asked for.

extension(name: 'str') -> 'str'

The file extension a format writes, including the leading dot.

describe() -> 'str'

Format list for --help, one per line.

write(name: 'str', source: 'Path', transcript: 'str', destination_dir: 'Path', **options) -> 'Path'

Write transcript in format name, returning the path written.

The destination filename is the source’s stem plus the format’s extension, which is what keeps a document’s outputs sitting together under one name.

Except when that would collide with the source itself. A PDF transcribed to a searchable PDF wants the same name as the original, and the batch pipeline moves the original into the same directory – so one silently overwrote the other, with the run reporting success either way. Such an output is written as <stem>.searchable<ext> instead. Only a source whose own extension is also an output format can hit this, which today means PDF alone.

Telemetry

tetrak_ocr.telemetry

Append-only run log: what was processed, by which backend, and how long it took.

A long batch is otherwise silent about its own cost. auto-local prints its score table only once every backend has finished, so a file that takes forty minutes looks identical to one that has hung – and when it does finish, the table says which backend won but not which one spent the time.

This module records that as it happens. Every event is one JSON object on one line, flushed immediately, so tail -f on the log is a live progress view and the finished file is a dataset:

{"ts": "...", "run": "20260822T115103Z-7f3a", "event": "backend_finish",
 "file": "camera-1919.pdf", "backend": "marker", "seconds": 2246.8, ...}

JSON Lines rather than CSV because the fields differ per event and will grow; rather than a database because appending a line is atomic enough for a sequential pipeline and needs nothing installed to read it back.

The sink is process-global. That is deliberate: backends expose a fixed ocr_image(path) -> str signature with nowhere to thread a logger through, and the fan-out is documented as sequential, so there is no concurrent writer to coordinate. When no run has been started every call is a cheap no-op, which is what keeps the library usable as a library.

TelemetryLog(path: 'Path', run_id: 'str') -> 'None'

One run’s worth of events, appended to a JSON Lines file.

start_run(path: 'Path | str | None' = None, **fields) -> 'TelemetryLog | None'

Begin recording, returning the log (or None when disabled).

finish_run(**fields) -> 'None'

Close out the current run and detach the sink.

record(event: 'str', **fields) -> 'None'

Emit an event if a run is active, otherwise do nothing.

active() -> 'bool'

No docstring. This function is public and undocumented.

log_path() -> 'Path | None'

No docstring. This function is public and undocumented.

describe_file(path: 'Path') -> 'dict'

The file facts worth having on every record about it.

Page count matters more than byte size here and is the thing people are surprised by: a 3 MB PDF that is ten pages is ten documents of work, and reading a per-file duration without it invites the wrong conclusion. It is read cheaply and never allowed to fail – a page count is not worth an exception in the middle of a batch.

timed(event: 'str', **fields)

Time a block and emit event with its duration, success or not.

Emits on the failure path too, because a backend that raised after twenty minutes is exactly the thing worth having in the log.

Batch pipeline

tetrak_ocr.batch

Batch OCR processor.

Scans the workspace/scans/ directory for supported image and PDF files, runs OCR on each one, writes the extracted text to a Markdown file in workspace/processed/, and moves the original file in alongside it.

After processing, each file pair looks like: workspace/processed/ my-postcard.jpg ← original moved here my-postcard.md ← extracted text

Usage (run from the repository root): tetrak-ocr batch tetrak-ocr batch –backend auto-local tetrak-ocr batch –backend claude tetrak-ocr batch –backend tesseract –contrast 3.0

build_ocr_fn(backend: str, contrast: float | None = None, psm: int | None = None, auto: bool = False, with_paddle_vl: bool = False) -> collections.abc.Callable[[pathlib.Path], str]

Return the OCR callable for the chosen backend.

Backend resolution lives in :mod:tetrak_ocr.registry — this used to duplicate it as an if/elif chain of imports, which meant a new backend had to be registered in two places and a missing optional dependency surfaced as a bare ImportError.

Only Tesseract takes tuning arguments and only auto-local takes --with-paddle-vl; the rest ignore them, and passing one is reported rather than silently dropped.

Both Tesseract names take the flags. tesseract-auto is the same backend with auto-configuration switched on by the registry, so it honours --contrast and --psm for whichever of the two you name; it used to drop them, because the wiring below tested == "tesseract" where the warning above tests startswith, leaving that one name in neither branch.

A value left as None is what lets auto-configuration fill it. analyse_image() supplies exactly the parameters the caller did not, so substituting a default here silently turns auto off – which is what made --auto inert on the tesseract backend: contrast and psm were always filled before the backend could look at the image.

supported_extensions(backend: str) -> set[str]

File extensions backend accepts.

auto-local routes PDFs to Tesseract, so it reports Tesseract’s full set — the registry reads this from each backend’s own SUPPORTED_EXTENSIONS.

transcribe_per_page(image_path: pathlib.Path, ocr_fn: collections.abc.Callable[[pathlib.Path], str]) -> list[str]

Transcribe image_path one page at a time, returning a text per page.

Needed only for a multi-page searchable PDF, which has to lay each page’s words onto the page they came from. It costs roughly twice a whole-document pass – measured, and not caused by anything we can avoid: reusing one Marker converter across the pages made no difference. See design research note 001 (Marker throughput and multi-page PDFs).

That is why the caller takes this route only when the pdf format is actually asked for on a multi-page document, rather than always.

process_file(image_path: pathlib.Path, ocr_fn: collections.abc.Callable[[pathlib.Path], str], triaged: list[str] | None = None, quality_gate: bool = False, backend: str | None = None, formats: list[str] | None = None, pdf_full_size: bool = False, pdf_font: str | None = None) -> None

OCR a single file, write Markdown output, and move it to workspace/processed/.

A transcript that scores below MIN_QUALITY_THRESHOLD goes to the triage queue instead of workspace/processed/, so a bad read is triaged rather than silently written out. There are two routes to that, because there are two things that know a transcript is poor:

The gate is opt-in for single backends rather than always on: it changes which files reach workspace/processed/, and turning it on silently would start diverting output for anyone already running --backend tesseract.

Args: image_path: Path to the file inside workspace/scans/. ocr_fn: The OCR callable to use. triaged: Optional list to append filenames routed to triage. quality_gate: Score this backend’s transcript and gate on it. backend: Backend name, used to label the triage manifest. formats: Output formats to write, defaulting to Markdown alone. Every requested format is written for the same transcript – the OCR runs once regardless of how many are asked for.

main(args=None) -> int

Parse arguments and process all supported files found in workspace/scans/.

Returns a process exit code – 0 on success, 1 if any file failed or workspace/scans/ is missing – so that callers can propagate a partial failure instead of reporting success. Every failure this function decides on is returned rather than raised, which is what lets cli.main honour its own -> int contract.

The one exception is not ours: argparse raises SystemExit on an unparseable argument list before this function regains control. Callers driving it programmatically with untrusted arguments should expect that.

Errors

tetrak_ocr.errors

Exceptions raised by the package.

These exist mainly so that a missing optional backend fails as a library should — with a catchable exception naming the extra to install — rather than by printing to stderr and calling sys.exit(1), which is what the original scripts did. That was reasonable when each file was run directly; it is not, when a caller imports one backend and wants to fall back to another.

OcrPipelineError(…)

Base class for every error this package raises.

MissingBackendError(backend: 'str', extra: 'str', packages: 'str') -> 'None'

A backend was requested but its optional dependency is not installed.

The message names the extra to install, because “No module named ‘paddleocr’” tells a user what broke but not what to do about it.

UnknownBackendError(name: 'str', known: 'list[str]') -> 'None'

A backend name was requested that does not exist.

UnknownFormatError(name: 'str', valid: 'list[str]') -> 'None'

An output format was requested that does not exist.

Carries the valid names: the useful thing to tell someone who typed --format pdff is what they could have typed instead.

MultiPageNotSupportedError(path: 'Path', backend: 'str', pages: 'int', use_instead: 'str') -> 'None'

A multi-frame image was given to a backend that reads only frame 0.

Raised instead of transcribing page one and returning it as though it were the whole document. Archival TIFF is often multi-page, and a partial transcript that looks complete is the failure mode this package treats as worst: it passes quality scoring, enters the archive, and misleads every reader after that.

MultiPagePdfError(path: 'Path', pages: 'int') -> 'None'

A searchable PDF was asked for on a document with more than one page.

The text layer carries the transcript the pipeline chose, and a backend returns one transcript for a whole document rather than one per page. There is no honest way to divide it: the pages would be guesses, and a reader searching a 21-page programme would be sent to the wrong leaf while the file looked entirely correct.

So this refuses, for the same reason :class:MultiPageNotSupportedError does. Per-page transcripts arrive with the positioned text layer – see brief 008 (searchable PDF).