Tetrak OCR
GitHub ↗
The project

About us

Two of us build Tetrak: one from data engineering, one from collections. The tool sits where those two meet.

That pairing is deliberate. The routing, the scoring and the triage queue are pipeline problems — the kind that turn up wherever unreliable input has to be made trustworthy. What counts as a usable transcript, and what a collection needs in order to be findable, is not a pipeline problem at all.

Pipelines and engineering

Stephen Masters

Stevie has spent close to three decades building software that turns messy, real-world data into something trustworthy; from foreign exchange pricing and compliance systems in banking, to power trading and generation platforms for energy companies, to the commodities data platforms he builds today. That work has given him a specific kind of expertise: designing pipelines that take in unreliable input, clean and validate it, and route each piece to wherever it’s handled best — the same problem Tetrak solves for OCR. He works across the stack, from cloud architecture down to the code itself, mostly in Python and data engineering these days, with Java, Angular and React still in the toolkit. He’s drawn to problems where messy signals (market data, sensor readings, satellite images) turn out to predict something useful, and Tetrak’s unruly archives are simply the latest iteration of that.

Collections systems and metadata

Yvette Mankerian

Yvette has spent over two decades running the systems that hold cultural institutions’ collections together — taxonomy and metadata, permissions, ingest and delivery, upgrades, integrations, and the documentation that keeps it all usable. That’s taken her through the Walt Disney Company, managing digital image assets and later building a taxonomy and metadata framework from scratch for Imagineering’s R&D collection; the Natural History Museum of Los Angeles County; and the Academy of Motion Picture Arts and Sciences, where she led the migration of the public collections catalogue to a new platform and now administers the systems behind roughly 2.6 million cataloguing records. Across all of it, she’s worked on both sides of the same problem: what a collection needs to be findable, and what a system needs to make that possible. That grounding — knowing exactly what curators and archivists deal with day to day, in real, unruly collections — is what Tetrak is built on.