npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

sttbench

v0.1.0

Published

Benchmark real-time speech-to-text providers on your own audio. Streaming replay, word error rate, and time-to-final-segment latency.

Readme

sttbench

Benchmark real-time speech-to-text providers on your own audio.

Published STT benchmarks tell you how providers perform on their test set — read audiobooks, curated podcasts, clean studio recordings. If you are building a voice agent, what you actually need to know is how they do on your calls: your accents, your background noise, your product vocabulary, your phone codec.

sttbench points at a folder of your recordings, replays each one through every provider at real time the way a live caller would, and reports three numbers side by side:

| | | |---|---| | Accuracy | Word error rate against your own reference transcripts | | Latency | Time from end-of-speech to final text (TTFS) — median and p95 | | Cost | What the run would cost per hour of audio, at the rate you set |

It runs on your machine, with your API keys. Nothing is uploaded anywhere except to the providers you name.

npx sttbench run ./calls

Quick start

1. Put your audio in a folder. Any format ffmpeg can read.

calls/
  0001.wav
  0002.mp3
  0003.m4a

2. Add reference transcripts — a .txt beside each file, same name. This is what makes accuracy measurable.

calls/
  0001.wav
  0001.txt      <- what was actually said
  0002.mp3
  0002.txt

Without references, sttbench still measures latency and cost, but word error rate is not reported rather than shown as zero. You cannot measure accuracy without ground truth, and pretending otherwise is how benchmarks mislead people.

Hand-transcribing audio is the real cost of benchmarking, and it is also the only way to get ground truth. A workable middle path is to draft with a batch model and then correct by ear:

OPENAI_API_KEY=... node scripts/draft-references.mjs ./calls
# writes call-01.draft.txt … review, fix, then:
mv calls/call-01.draft.txt calls/call-01.txt

sttbench only reads .txt, so nothing is scored until you have deliberately promoted a file.

🚨 Never draft with a provider you are benchmarking. If Deepgram writes the reference, Deepgram scores ~0% WER by construction and the benchmark measures nothing. The script refuses to run against the benchmarked set for exactly this reason.

Even a neutral batch model shares training data and failure modes with the systems under test, which is why the human pass is not optional. Concentrate on the words a benchmark exists to measure — names, numbers, times, dish names, addresses. Fixing "Geetha" that was drafted as "Gita" is worth more than fixing ten commas.

3. Set the API keys for the providers you want.

export SONIOX_API_KEY=...
export DEEPGRAM_API_KEY=...
export ASSEMBLYAI_API_KEY=...

npx sttbench providers      # shows which are ready

4. Run it.

npx sttbench run ./calls --language en
provider    model                          WER     ±  lag p50  lag p95  cost/hr    total  ok/fail
─────────────────────────────────────────────────────────────────────────────────────────────────
deepgram    nova-3                        1.0%  1.6%    940ms   1893ms  $0.4600  $0.0039      3/0
soniox      stt-rt-v4                     2.0%  3.1%    634ms    696ms  $0.1200  $0.0010      3/0
assemblyai  universal-streaming-english  11.0%  3.6%    515ms    515ms  $0.1500  $0.0013      3/0

You also get results.json (every transcript, every metric) and a self-contained report.html you can share.


Why real-time replay matters

Most benchmarks upload a file and time the response. That measures a batch API. Streaming providers behave completely differently when audio arrives at the speed a person talks: endpoint detection fires on silence, interim results get revised, and the finalization delay that decides how fast your agent can answer only exists in that regime.

sttbench streams 100ms chunks on a real-time clock — so a 10-minute recording takes about 10 minutes, and the latency numbers mean something.

The latency number is not "time to first transcript"

That metric is dominated by when the speaker started talking. On a recording with four seconds of ringing at the front, every provider scores ~4000ms and you have measured the recording, not the vendor.

sttbench measures finalization lag: how long after the speech ended the final text arrived.

lag = wall-clock when the final arrived − the segment's end time in the audio

Both p50 and p95 are reported, because a voice agent is judged on its worst turns — a provider with a good median and a heavy tail feels broken in conversation.


The LLM judge

WER treats every word equally. Mishearing "um" and mishearing "party of eight" cost the same, and they are not the same mistake.

So sttbench also runs an LLM judge scoring accuracy, entity correctness, segmentation and hallucination 0–10. It runs automatically whenever ANTHROPIC_API_KEY or OPENAI_API_KEY is set — one run gives you the whole picture, and the header always tells you whether it ran:

  judge: anthropic claude-sonnet-5
npx sttbench run ./calls --domain "restaurant phone bookings"   # judge on if a key exists
npx sttbench run ./calls --no-judge                             # skip it
npx sttbench run ./calls --judge openai                         # force a vendor

The judge never sees vendor names — transcripts are relabelled A/B/C and shuffled, seeded per sample, so the model cannot apply a brand prior and position bias averages out across your corpus.

It is deliberately secondary. WER leads because anyone can recompute it from the same inputs. A benchmark whose headline number cannot be independently reproduced is an opinion.

The two can disagree, and that is the point — a provider can have a slightly higher WER while making errors that matter less.


Options

sttbench run <dir-or-file> [options]

  -p, --providers <ids>     Comma list (default: every provider with its key set)
      --model <pairs>       Per-provider model, e.g. soniox=stt-rt-v4,deepgram=nova-3
      --cost <pairs>        Override $/hour, e.g. soniox=0.10
  -l, --language <code>     Language hint, e.g. en, en-AU, multi
      --limit <n>           Only benchmark the first N audio files
      --concurrency <n>     Parallel streams (default 4)
      --sample-rate <hz>    PCM sample rate (default 16000)
      --rank <axis>         wer | latency | cost (default wer)
      --domain <text>       Domain hint for the judge
      --judge <vendor>      Force the judge vendor: anthropic | openai
      --no-judge            Skip the LLM judge
      --judge-model <id>    Override the judge model
  -o, --out <dir>           Output directory (default sttbench-results)
      --cache-dir <dir>     Decoded PCM cache (default .sttbench-cache)
      --json                Only write files; no table
      --keep-numbers        Don't normalize spelled numbers before WER
      --keep-fillers        Count "um"/"uh" as words
      --keep-contractions   Don't expand contractions

sttbench providers          List providers and whether their keys are set

Supported providers

| id | default model | key | |---|---|---| | soniox | stt-rt-v4 | SONIOX_API_KEY | | deepgram | nova-3 | DEEPGRAM_API_KEY | | assemblyai | universal-streaming-english | ASSEMBLYAI_API_KEY |

Adding one is a single file plus a registry entry — see docs/providers.md. PRs welcome.

Requirements

  • Node.js 20+
  • ffmpeg on your PATH (brew install ffmpeg / apt install ffmpeg)

Use it as a library

import { computeWer, runBenchmark, discoverSamples } from 'sttbench';

const samples = await discoverSamples({ input: './calls' });
const report = await runBenchmark({ samples, providers, /* … */ });

How to read the results honestly

  • A small corpus proves nothing. With five recordings, a 1% WER gap is noise. Read the ± spread column; sttbench prints a warning below that threshold.
  • Cost is a calculation, not an invoice. It multiplies audio duration by a published list price. Your contract almost certainly differs — override it with --cost.
  • There is no single winner. sttbench will not compute a blended score. Accuracy, latency and cost trade off against each other, and which one dominates depends on what you are building. Any weighting we picked would be an opinion dressed up as a measurement.
  • Your reference transcripts are the ceiling. If they were produced by an ASR system rather than a human, you are partly measuring agreement with that system.

The full set of rules — every normalization step, every aggregation choice and why — is in docs/methodology.md.

Contributing

See CONTRIBUTING.md. New providers are the most useful contribution, and the provider interface is deliberately small.

Who made this

Built at Forktime, where we run STT on restaurant phone calls and needed to answer this question for ourselves.

Disclosure: we are a paying customer of some providers benchmarked here. That is exactly why the judge is blind to vendor identity, why WER is the headline number, and why every run ships the raw transcripts and a reproduce command. Do not take our word for the result — run it on your own audio.

License

MIT