n8n-nodes-typecast
v0.2.0
Published
n8n community node for Typecast API - AI text to speech with emotion control
Maintainers
Readme
n8n-nodes-typecast
An n8n community node for Typecast, the AI voice synthesis API behind Neosapience's SSFM speech models.
Turn text into a finished audio file inside a workflow - pick a voice from a searchable list, choose how the emotion is decided, and hand the result straight to Google Drive, S3, Slack or Telegram as binary data.
n8n is a fair-code licensed workflow automation platform.
Installation · Credentials · Operations · Choosing a voice · Emotion control · API reference · Output · Development
Installation
In n8n, go to Settings -> Community Nodes -> Install and enter:
n8n-nodes-typecastFor a self-hosted instance you can also install it directly:
npm install n8n-nodes-typecastCredentials
- Open the Typecast developer dashboard and copy your API Key.
- In n8n, create a Typecast API credential and paste the key.
- Hit Test - it calls
GET /v2/voicesand confirms the key works before you build anything.
The key travels as an X-API-KEY header on every request.
Operations
Speech
| Operation | Endpoint | Notes |
| --- | --- | --- |
| Text to Speech | POST /v1/text-to-speech | Returns the audio as binary data, not a URL |
Voice
| Operation | Endpoint | Notes |
| --- | --- | --- |
| Get | GET /v2/voices/{voice_id} | One voice by ID |
| Get Many | GET /v2/voices | Filter by model, gender, age and use case |
Choosing a voice
Voice is a resource locator with two modes.
From List queries /v2/voices filtered by the model you selected, so the dropdown only offers voices that model can actually speak with. The search box filters on name, ID, gender, age and use case, and results are sorted by name. Each entry shows the voice ID and its supported models in the description line.
ID takes a raw voice ID such as tc_60e5426de8b95f1d3000d7b5. Built-in voices start with tc_.
Picking a voice the selected model does not support fails at synthesis time, not at selection time, which is why the list is narrowed by model rather than showing the whole catalogue.
Emotion control
Typecast decides emotion differently per model, and the request shape differs with it. The node builds the right one for you; this table is what it does under the hood.
| Model | Mode | Fields sent |
| --- | --- | --- |
| ssfm-v30 | Smart | emotion_type: "smart" plus optional previous_text / next_text |
| ssfm-v30 | Preset | emotion_type: "preset" plus emotion_preset |
| ssfm-v21 | (only mode) | emotion_preset alone - this model rejects emotion_type |
Smart infers the emotion from surrounding context. Fill Previous Text and Next Text under Additional Options with the lines that come before and after, and the model reads the mood from them. Both are ignored outside Smart mode, because the API's preset variant does not accept them.
Preset picks the emotion by hand. ssfm-v30 offers seven presets - normal, happy, sad, angry, whisper, toneup, tonedown. ssfm-v21 offers four - normal, happy, sad, angry - and the node shows only those when that model is selected.
API reference
What the node sends to POST /v1/text-to-speech, and which n8n field controls each part.
| Field | n8n field | Constraint |
| --- | --- | --- |
| voice_id | Voice | tc_ for built-in voices, uc_ for cloned ones |
| text | Text | 1 to 2000 characters |
| model | Model | ssfm-v30 or ssfm-v21 |
| language | Language | ISO 639-3, case-insensitive. Omitted for auto-detect |
| prompt | Emotion Type / Emotion Preset | See Emotion control |
| output.audio_format | Additional Options -> Audio Format | wav or mp3, default wav |
| output.audio_pitch | Additional Options -> Audio Pitch | -12 to +12 semitones |
| output.audio_tempo | Additional Options -> Audio Tempo | 0.5 to 2.0 |
| output.volume | Additional Options -> Volume | 0 to 200, default 100 |
| seed | Additional Options -> Seed | uint32, 0 to 4294967295 |
Language coverage differs by model: ssfm-v30 handles 37 languages, ssfm-v21 handles 27. The dropdown lists the full v30 set.
target_lufs is accepted by the API as an alternative to volume but is not exposed by this node. The two are mutually exclusive at the server, so exposing both would need a validation rule that only earns its keep once someone wants loudness normalisation.
Output
The audio arrives as binary data in the field named data, ready for a file node without an HTTP Request node in between. The JSON side carries what produced it:
{
"voice_id": "tc_60e5426de8b95f1d3000d7b5",
"text": "Hello, this is a test.",
"model": "ssfm-v30",
"format": "wav"
}Items carry pairedItem, so item lineage survives into downstream nodes.
Failures surface the API's own message and error code rather than a byte dump. The audio endpoint answers as an array buffer, so an error body would otherwise be rendered as raw bytes; the node decodes it and rethrows with the status code attached.
Compatibility
Requires Node.js 20.15 or newer. Built against the n8n community node API version 1 and n8n-workflow 1.120.
Development
pnpm install
pnpm build # compile TypeScript and copy the icon
pnpm lint # n8n community node rules
pnpm test # compile, then run the unit testspnpm test asserts the prompt object built for each model and emotion mode. It needs no API key. This is the piece worth guarding: the wrong shape is silent in TypeScript and a hard 4xx at the API.
To try it in a local n8n instance, build it and symlink the package where n8n looks for custom nodes, then restart n8n:
pnpm build
mkdir -p ~/.n8n/custom/node_modules
ln -sfn "$PWD" ~/.n8n/custom/node_modules/n8n-nodes-typecastNodes loaded this way register under the package name CUSTOM, so a workflow JSON must reference the type as CUSTOM.typecast rather than n8n-nodes-typecast.typecast. Installing through Settings -> Community Nodes uses the full package name instead.
Upgrading to 0.2.0
Existing 0.1.x workflows keep working. Every parameter kept its name, so saved values are read back unchanged.
Fixed in 0.2.0:
ssfm-v21works again. The node sentemotion_typefor every model, but the API'spromptis a discriminated union andssfm-v21takes a legacy shape that rejects that field, so every v21 request failed. The prompt is now assembled per model.- Previous / Next Text no longer leak into preset requests. The preset variant of the union does not accept them.
ssfm-v21no longer offers presets it cannot speak. It accepts four; the UI listed all seven of v30's.- The voice search box works. The loader declared
searchable: truewhile ignoring the filter n8n passes in, so typing changed nothing. - The voice list is narrowed by model instead of returning the whole catalogue, which made it hard to scroll and let you pick a voice the model could not use.
- The voice loader no longer throws on a wrapped response body or a voice with no
modelsfield. - API errors are readable, instead of the JSON error body being rendered as an array buffer.
seed: 0is sent. A valid uint32 was being dropped as falsy.- Output items now carry
voice_id,text,modeland apairedItem.
Resources
Typecast publishes an official node as @neosapience/n8n-nodes-typecast, which additionally covers voice cloning, streaming and word-level timestamps.
