Model Evaluation Web App
A dual-mode evaluation web app that benchmarks speech-to-text and translation models on a single Spanish-language clinical session — reference-free and reference-based metrics, an LLM judge audited by human scoring, and a fully offline demo. Built solo with AI-assisted development.
/ The Challenge
Picking a speech-to-text or translation model for a clinical setting is usually done on vibes — you trust a provider's marketing and hope. This project set out to make that decision measurable: benchmark a field of ASR and translation models on the same real-world Spanish therapy audio across accuracy, cost, and speed, and present the result so a non-engineer could read it, not just run it. Two constraints shaped everything. It had to serve two audiences from one codebase — a transcription-engine decision for a Spanish-language mental-health app, and an AI speech-interpreting accuracy study — without forking into two tools. And it had to stay cheap and private: a few dollars of total API spend, no real patient recording ever touched, and a version anyone could open and explore with no API keys at all.
/ Approach
- 01
One schema, every stage. Transcription, translation, and a native audio→text lane are modeled as composable stages that all emit the same Record. Cost math, the metric registry, the report, and the UI are each written once against that shape — so adding a provider is a lazily-imported adapter file, not a new pipeline.
- 02
A metric registry that turns on per stage. Reference-free signals (inter-model disagreement, self-consistency, an LLM rubric judge, GEMBA-MQM) run on every output; reference-based scores switch on automatically when a human "gold" file is attached — WER/CER for transcription, BLEU/chrF++/TER for translation — each applied only to the stage it can actually measure.
- 03
An LLM judge, audited by a human. A Claude judge scores every output 1–5 on a clinical rubric (fidelity, omissions, clinical-term accuracy, register). Because a grader with no ground truth is just an opinion, a human sample — scored on the highest-disagreement chunks first — sits beside it to validate whether the judge can be trusted on the rest.
- 04
Offline demo, live runs gated. The whole UI renders from committed, synthetic cached runs with zero API calls, so it demos safely and reproducibly. Real runs are guarded by a development slice (~45 seconds of audio) and a single-digit-dollar budget cap, with a build-time-verified pricing table behind every cost shown.
- 05
A result you can read, not just run. Wrapped in a themeable (light/dark) design system with collapsible sections, side-by-side A/B run comparison, ranked accuracy tables with plain-English metric explanations, and a one-click, self-contained HTML report.
/ Selected Technical Challenges
-
A judge that costs a rounding error. The Claude judge originally ran with adaptive thinking on, which quietly made it the single largest cost in a run. Since the judge only needs to emit a small structured JSON verdict, disabling thinking cut its cost dramatically with no loss in grading quality — and the judge model, coverage, and sampling are all tunable, so the audit can be run for cents.
-
Metrics that don't lie about cost — or hang. Every price in the app is verified against the provider's published rate at build time, so the on-screen cost estimate is real, not a guess. One metric, document-level TER, turned out to be roughly quadratic and could stall a re-scoring pass for the better part of an hour; it was made opt-in so it corroborates BLEU/chrF++ without ever blocking a run.
-
Offline-first for a live-API tool. The hardest product constraint was that a tool whose entire job is calling paid speech APIs still had to open, render, and be fully explorable with no keys, no network, and no spend. Cached runs are committed as first-class data (audio excluded), the metric and cost layers read those records identically to a live run, and the demo is the same code path as production — just fed from disk.
/ Outcomes
-
One codebase serves two deliverables — a clinical ASR-engine decision and an AI-interpreting accuracy study — with no forked code.
-
80+ models across five providers (OpenAI, Anthropic, Google, Deepgram, AssemblyAI) inventoried; any wired model runs through the same pipeline and scoring.
-
A full evaluation of the reference run cost $2.67 total, every number backed by a build-time-verified pricing table.
-
Reference-based metrics evidence the translation stage: claude-sonnet-5 (literal) led at BLEU 42.7 / chrF++ 72.2, and the native audio→English lane came within ~1 BLEU of the best cascade — a concrete cascaded-vs-end-to-end data point.
-
An LLM judge validated against a human baseline, plus side-by-side A/B run comparison and a downloadable, self-contained HTML report.
-
Fully offline demo — the entire UI renders from cached runs with zero keys and zero spend.