@amirhosseinnouri/voxgen
v1.0.0
Published
Turn a text file into a narrated audio file with Fish Audio.
Maintainers
Readme
voxgen
Turn a text file into a narrated audio file. Fish Audio does the speaking; this tool does the parts around it — chunking, cost, caching, and joining the pieces back into one clean recording.
The default backend is s2.1-pro-free, which is the same model as s2.1-pro at $0. A
plain run costs nothing.
Setup
bunx @amirhosseinnouri/voxgen ./article.mdOr from source:
git clone https://github.com/amirhosseinNouri/voxgen
cd voxgen
bun installPut your key in .env:
FISH_API_KEY=... # https://fish.audio/go-api/api-keysffmpeg is needed only to write something other than WAV — mp3, opus, flac and m4a all go
through it.
Setting up with AI
To install, configure and first-run voxgen with an AI assistant, tell it to use the
setup skill: use the setup skill to set up voxgen. It walks through picking how to run
it, getting a key into .env, deciding on ffmpeg, verifying with a free run, and choosing
a voice — in that order, verifying each step as it goes.
Usage
bun start ./article.md # → output/article-<timestamp>/audio.mp3
bun start "Good morning." # text that is not a file is read literally
cat notes.md | bun start --stdin # or pipe it in
bun start ./article.md -o talk.opus # write exactly here, in this container
bun start voices calm # find a voice id for --voiceOptions:
| Flag | Meaning |
| --- | --- |
| -o, --output <path> | Write here instead of output/<name>-<timestamp>/ |
| --format <fmt> | wav, mp3, opus, flac, m4a (default mp3) |
| --voice <id> | A voice id from voxgen voices |
| --model <id> | Fish backend (default s2.1-pro-free) |
| --speed <n> | 0.5–2.0; 1 is the voice's own pace |
| -y, --yes | Skip the cost confirmation |
Each run without -o writes a fresh output/<name>-<timestamp>/ containing audio.<fmt>
and script.txt — the normalized text that was actually spoken, which is what to check
when a word comes out wrong. output/ is git-ignored.
How it works
- Normalization — Markdown markers are flattened first, because a narrator should read
"Title", not "hash Title". Headings, bullets, emphasis and inline code lose their
punctuation; a link reads as its label rather than its URL. A
#inside a sentence survives, since only leading heading markers are stripped. - Voice pinning — every request carries an explicit
reference_id. Fish accepts a request without one and quietly picks a speaker for it, which is harmless for a single sentence and ruinous for a document: an article synthesized in 84 requests comes back in 84 different voices. There is therefore always a voice, defaulting toSarah; change it with--voice,FISH_VOICE_ID, orvoxgen voices. - Chunking — the script is cut into pieces of about 1500 UTF-8 bytes. Paragraphs are the outer unit and never share a chunk; within a paragraph the split is by sentence, and only by word when one sentence is over budget on its own. A word longer than the budget — a URL, a hash — is cut on a character boundary rather than dropped, and never in the middle of a multi-byte character. Chunks exist for progress, caching and retry granularity, not for a provider limit.
- Cost estimate — Fish bills per UTF-8 byte of input, so the entire bill is knowable before a single byte goes out. Chunks already in the cache are subtracted and the run prints what is left to send, with its price, before asking to continue. Nothing reaches the provider until that prompt is answered — and on the free backend there is no prompt, because there is nothing to decide.
- Synthesis — each chunk comes back as raw 16-bit mono PCM rather than mp3. Raw samples join gaplessly, where two mp3 streams glued together leave an encoder-padding click at every seam. Each chunk is written to the cache before the next request goes out, so an interrupted run resumes where it stopped and never pays for the same sentence twice. Transient failures — 429, 5xx — are retried with backoff; a rejected request is not, because sending the same bytes again buys the same rejection. A long article is a lot of small requests, so they go out in a sliding window of four rather than one at a time — but they are written strictly in order, so the recording is identical to what a sequential run would produce.
- Joining — the chunks are streamed straight into a WAV file, with a short silence inserted wherever the source had a blank line, so paragraphs land as pauses instead of the voice running two thoughts together. The file is written as it goes and its header patched at the end, so an hour of audio never has to fit in memory. A run that fails part-way deletes its output: a truncated WAV is still a playable WAV, and one left on disk looks exactly like a finished result until someone reaches the end of it.
- Encoding — anything but WAV is transcoded by ffmpeg at rates suited to speech (96k mp3, 48k opus), and the intermediate is removed once that succeeds.
Caching
Synthesized audio is cached under .cache/voxgen/, keyed by a hash of the request: the
text, the backend, the voice, the speed, the volume and the sample rate. Editing one
paragraph of a long article re-synthesizes that paragraph and nothing else. Changing the
voice correctly invalidates everything. Cache files are written under a temporary name and
renamed into place, so a run killed mid-write cannot leave a half-length chunk that the
next run treats as complete.
Delete .cache/voxgen/ to start paying again.
Configuration
Everything has a default; nothing but the key is required.
| Variable | Default | |
| --- | --- | --- |
| FISH_API_KEY | — | Required |
| FISH_MODEL | s2.1-pro-free | Backend, sent as the model header |
| FISH_VOICE_ID | 9335…406a (Sarah) | Voice; same values as --voice. Never empty — see step 2 |
| FISH_BASE_URL | https://api.fish.audio | |
| TTS_LATENCY | normal | balanced trades quality for first-byte latency |
| TTS_SPEED | 1 | 0.5–2.0 |
| TTS_VOLUME | 0 | dB adjustment applied by the provider |
| TTS_SAMPLE_RATE | 44100 | |
| TTS_PRICE_PER_MILLION_BYTES | published rate | Override for the cost estimate |
| CHUNK_BYTES | 1500 | Request size, in UTF-8 bytes |
| PARAGRAPH_PAUSE_MS | 350 | Silence at a blank line |
| CONCURRENCY | 4 | Requests in flight at once; Fish's entry tier allows 5 |
| CACHE_DIR | .cache/voxgen | |
Prices are per UTF-8 byte, not per character. One Persian or CJK character is three
bytes, so a document that looks half the length of an English one can cost three times as
much — voxgen counts and prices in bytes for exactly that reason.
Anything invalid is rejected at startup with the variable named, rather than surfacing mid-run once part of the script has already been paid for.
Development
bun test # 118 tests, no network
bun run typecheck