sttbench
v0.1.0
Published
Benchmark real-time speech-to-text providers on your own audio. Streaming replay, word error rate, and time-to-final-segment latency.
Maintainers
Readme
sttbench
Benchmark real-time speech-to-text providers on your own audio.
Published STT benchmarks tell you how providers perform on their test set — read audiobooks, curated podcasts, clean studio recordings. If you are building a voice agent, what you actually need to know is how they do on your calls: your accents, your background noise, your product vocabulary, your phone codec.
sttbench points at a folder of your recordings, replays each one through every provider at real time the way a live caller would, and reports three numbers side by side:
| | | |---|---| | Accuracy | Word error rate against your own reference transcripts | | Latency | Time from end-of-speech to final text (TTFS) — median and p95 | | Cost | What the run would cost per hour of audio, at the rate you set |
It runs on your machine, with your API keys. Nothing is uploaded anywhere except to the providers you name.
npx sttbench run ./callsQuick start
1. Put your audio in a folder. Any format ffmpeg can read.
calls/
0001.wav
0002.mp3
0003.m4a2. Add reference transcripts — a .txt beside each file, same name. This
is what makes accuracy measurable.
calls/
0001.wav
0001.txt <- what was actually said
0002.mp3
0002.txtWithout references, sttbench still measures latency and cost, but word error rate is not reported rather than shown as zero. You cannot measure accuracy without ground truth, and pretending otherwise is how benchmarks mislead people.
Hand-transcribing audio is the real cost of benchmarking, and it is also the only way to get ground truth. A workable middle path is to draft with a batch model and then correct by ear:
OPENAI_API_KEY=... node scripts/draft-references.mjs ./calls
# writes call-01.draft.txt … review, fix, then:
mv calls/call-01.draft.txt calls/call-01.txtsttbench only reads .txt, so nothing is scored until you have
deliberately promoted a file.
🚨 Never draft with a provider you are benchmarking. If Deepgram writes the reference, Deepgram scores ~0% WER by construction and the benchmark measures nothing. The script refuses to run against the benchmarked set for exactly this reason.
Even a neutral batch model shares training data and failure modes with the systems under test, which is why the human pass is not optional. Concentrate on the words a benchmark exists to measure — names, numbers, times, dish names, addresses. Fixing "Geetha" that was drafted as "Gita" is worth more than fixing ten commas.
3. Set the API keys for the providers you want.
export SONIOX_API_KEY=...
export DEEPGRAM_API_KEY=...
export ASSEMBLYAI_API_KEY=...
npx sttbench providers # shows which are ready4. Run it.
npx sttbench run ./calls --language enprovider model WER ± lag p50 lag p95 cost/hr total ok/fail
─────────────────────────────────────────────────────────────────────────────────────────────────
deepgram nova-3 1.0% 1.6% 940ms 1893ms $0.4600 $0.0039 3/0
soniox stt-rt-v4 2.0% 3.1% 634ms 696ms $0.1200 $0.0010 3/0
assemblyai universal-streaming-english 11.0% 3.6% 515ms 515ms $0.1500 $0.0013 3/0You also get results.json (every transcript, every metric) and a
self-contained report.html you can share.
Why real-time replay matters
Most benchmarks upload a file and time the response. That measures a batch API. Streaming providers behave completely differently when audio arrives at the speed a person talks: endpoint detection fires on silence, interim results get revised, and the finalization delay that decides how fast your agent can answer only exists in that regime.
sttbench streams 100ms chunks on a real-time clock — so a 10-minute recording takes about 10 minutes, and the latency numbers mean something.
The latency number is not "time to first transcript"
That metric is dominated by when the speaker started talking. On a recording with four seconds of ringing at the front, every provider scores ~4000ms and you have measured the recording, not the vendor.
sttbench measures finalization lag: how long after the speech ended the final text arrived.
lag = wall-clock when the final arrived − the segment's end time in the audioBoth p50 and p95 are reported, because a voice agent is judged on its worst turns — a provider with a good median and a heavy tail feels broken in conversation.
The LLM judge
WER treats every word equally. Mishearing "um" and mishearing "party of eight" cost the same, and they are not the same mistake.
So sttbench also runs an LLM judge scoring accuracy, entity correctness,
segmentation and hallucination 0–10. It runs automatically whenever
ANTHROPIC_API_KEY or OPENAI_API_KEY is set — one run gives you the whole
picture, and the header always tells you whether it ran:
judge: anthropic claude-sonnet-5npx sttbench run ./calls --domain "restaurant phone bookings" # judge on if a key exists
npx sttbench run ./calls --no-judge # skip it
npx sttbench run ./calls --judge openai # force a vendorThe judge never sees vendor names — transcripts are relabelled A/B/C and shuffled, seeded per sample, so the model cannot apply a brand prior and position bias averages out across your corpus.
It is deliberately secondary. WER leads because anyone can recompute it from the same inputs. A benchmark whose headline number cannot be independently reproduced is an opinion.
The two can disagree, and that is the point — a provider can have a slightly higher WER while making errors that matter less.
Options
sttbench run <dir-or-file> [options]
-p, --providers <ids> Comma list (default: every provider with its key set)
--model <pairs> Per-provider model, e.g. soniox=stt-rt-v4,deepgram=nova-3
--cost <pairs> Override $/hour, e.g. soniox=0.10
-l, --language <code> Language hint, e.g. en, en-AU, multi
--limit <n> Only benchmark the first N audio files
--concurrency <n> Parallel streams (default 4)
--sample-rate <hz> PCM sample rate (default 16000)
--rank <axis> wer | latency | cost (default wer)
--domain <text> Domain hint for the judge
--judge <vendor> Force the judge vendor: anthropic | openai
--no-judge Skip the LLM judge
--judge-model <id> Override the judge model
-o, --out <dir> Output directory (default sttbench-results)
--cache-dir <dir> Decoded PCM cache (default .sttbench-cache)
--json Only write files; no table
--keep-numbers Don't normalize spelled numbers before WER
--keep-fillers Count "um"/"uh" as words
--keep-contractions Don't expand contractions
sttbench providers List providers and whether their keys are setSupported providers
| id | default model | key |
|---|---|---|
| soniox | stt-rt-v4 | SONIOX_API_KEY |
| deepgram | nova-3 | DEEPGRAM_API_KEY |
| assemblyai | universal-streaming-english | ASSEMBLYAI_API_KEY |
Adding one is a single file plus a registry entry — see docs/providers.md. PRs welcome.
Requirements
- Node.js 20+
ffmpegon your PATH (brew install ffmpeg/apt install ffmpeg)
Use it as a library
import { computeWer, runBenchmark, discoverSamples } from 'sttbench';
const samples = await discoverSamples({ input: './calls' });
const report = await runBenchmark({ samples, providers, /* … */ });How to read the results honestly
- A small corpus proves nothing. With five recordings, a 1% WER gap is
noise. Read the
±spread column; sttbench prints a warning below that threshold. - Cost is a calculation, not an invoice. It multiplies audio duration by a
published list price. Your contract almost certainly differs — override it
with
--cost. - There is no single winner. sttbench will not compute a blended score. Accuracy, latency and cost trade off against each other, and which one dominates depends on what you are building. Any weighting we picked would be an opinion dressed up as a measurement.
- Your reference transcripts are the ceiling. If they were produced by an ASR system rather than a human, you are partly measuring agreement with that system.
The full set of rules — every normalization step, every aggregation choice and why — is in docs/methodology.md.
Contributing
See CONTRIBUTING.md. New providers are the most useful contribution, and the provider interface is deliberately small.
Who made this
Built at Forktime, where we run STT on restaurant phone calls and needed to answer this question for ourselves.
Disclosure: we are a paying customer of some providers benchmarked here. That is exactly why the judge is blind to vendor identity, why WER is the headline number, and why every run ships the raw transcripts and a reproduce command. Do not take our word for the result — run it on your own audio.
License
MIT
