unmodel
v0.5.1
Published
Validation layer for LLM API calls: catalog-aware, zod-powered request validation and response sanity checks for every provider — bring your own SDK or fetch.
Maintainers
Readme
unmodel checks your request against what the provider actually accepts, before you send it. Optional cross-provider params compile to real wire bodies. Raw responses get checked for truncation, refusals, filtering, usage, and cost. It never sends anything, never sees a credential. You keep fetch, your SDK, and your keys.
Every ❌/✅ in this README is pasted from a real tsc run or backed by a test — including the hero above, where both compile errors come from tsc against [email protected] and this package.
The hero is the whole pitch in one union. The official SDK types Sora's size the same for every model, and the one union is wrong in both directions:
// [email protected]
await client.videos.create({ model: "sora-2-pro", prompt: "…", size: "1920x1080" });
// ❌ error TS2322: Type '"1920x1080"' is not assignable to type 'VideoSize | undefined'.
// The 1080p that sora-2-pro documents, prices, and renders — the SDK won't compile it.
// (16- and 20-second clips? Also documented, also refused: VideoSeconds stops at "12".)
await client.videos.create({ model: "sora-2", prompt: "…", size: "1792x1024" });
// ✅ compiles — for a resolution sora-2 does not render (720p only), so the API refuses it
// unmodel — the documented matrix, per model
import { video } from "unmodel/openai";
video({ model: "sora-2-pro", prompt: "…", size: "1920x1080", seconds: "16" }); // ✅
video({ model: "sora-2", prompt: "…", size: "1792x1024" });
// ❌ error TS2322: Type '"1792x1024"' is not assignable to type 'SoraBaseSize | undefined'.
// sora-2 renders 720p only — the mistake the SDK compiles is a compile error here✨ Features
- 🎯 Per-model types —
size: "1792x1024"onsora-2is a compile error, not a 400 — and the 1080p the official SDK won't compile at all is justsora-2-pro's documented size - 🔤 Real autocomplete — the 23 real
gpt-image-2sizes, all 30 Gemini TTS voices; every list proven by a test - 🌐 One vocabulary, fourteen surfaces — chat, TTS, STT, image, image edit, video, lipsync, avatar, upscale, 3D, music, voice clone, voice design, realtime config
- 🔁
.toApi(provider)— move a validated chat request to another host serving the same model - 💸 Cost gates —
.safe({ maxCostUSD })blocks a runaway request before it leaves the process - 🧾 Response checks — truncation, refusals, filtering, usage, and catalog-priced cost from raw payloads
- 🗂️ Generated models.dev catalog — capabilities, context limits, pricing, deprecations
- 🪶 Zero SDK dependency — types-only entries emit no JavaScript; values entries cost ~1 KiB each
- ⚡ Runs on Node 20+, Bun, and Cloudflare Workers
npm install unmodel
# or: bun add unmodel🚀 Quick start
Write one request, validate it against the selected provider, send the body yourself:
import { toRequestInit } from "unmodel";
import { chat } from "unmodel/chat";
import { checkChat } from "unmodel/openai";
const request = chat({
model: "openai/gpt-5.2",
messages: [{ role: "user", content: "Explain this code in one sentence." }],
maxOutputTokens: 256,
});
JSON.stringify(request);
// → {"model":"gpt-5.2","messages":[...],"max_completion_tokens":256}
const { url, ...init } = toRequestInit(request);
const response = await fetch(url, {
...init,
headers: {
...init.headers,
authorization: `Bearer ${process.env.OPENAI_API_KEY ?? ""}`,
},
});
const report = checkChat(await response.json());
report.finishReason; // "stop", "length", ...
report.costUSD; // actual catalog-priced usage, when availableThe result is the wire body; toRequestInit packs URL, method, and static headers for fetch. Auth stays yours — no unmodel export takes a key.
🧭 Choose a surface
| Goal | Use | Input |
| --- | --- | --- |
| One shape across providers | unmodel/chat, unmodel/tts, unmodel/stt, etc. | camelCase params + "provider/model" |
| Exact provider API | unmodel/openai, unmodel/anthropic, etc. | provider-native fields + bare model id |
| Small cross-provider bundle | createChat from unmodel/chat/factory; media factories from their category entries | serves only registered providers |
| Move validated chat to another host | .toApi(provider) | an existing validated request |
| Types with no runtime | unmodel/<provider>/types, unmodel/types | nothing (the entries emit no JavaScript) |
| Runtime lists for pickers | unmodel/<provider>/values, unmodel/values | nothing (arrays out, ~1 KiB per import) |
Unified calls compile to provider-native params and finish in that provider's validator:
canonical params → provider wire params → provider validator → fetch or SDKProvider validators take provider-native fields, and the validated result is the exact wire body. Unified model refs split on the first slash: openrouter/anthropic/claude-opus-5 means provider openrouter, model anthropic/claude-opus-5.
🎨 The fifteen surfaces
| Task | Portable import | Provider-native example |
| --- | --- | --- |
| 💬 Chat | unmodel/chat | unmodel/openai, unmodel/anthropic, unmodel/google |
| 🗣️ Text to speech | unmodel/tts | unmodel/openai, unmodel/elevenlabs, unmodel/deepgram |
| 🎧 Speech to text | unmodel/stt | unmodel/openai, unmodel/deepgram, unmodel/assemblyai |
| 🖼️ Image generation | unmodel/image | unmodel/openai, unmodel/google, unmodel/black-forest-labs |
| ✏️ Image editing | unmodel/image-edit | unmodel/openai, unmodel/black-forest-labs, unmodel/ideogram |
| 🎬 Video generation | unmodel/video | unmodel/openai, unmodel/google, unmodel/runway |
| 👄 Lipsync | unmodel/lipsync | unmodel/fal, unmodel/heygen, unmodel/sync, unmodel/veed |
| 🧑🎤 Avatar | unmodel/avatar | unmodel/fal, unmodel/heygen, unmodel/sync, unmodel/veed |
| 🔍 Upscale | unmodel/upscale | unmodel/fal, unmodel/topaz |
| 🧊 3D generation | unmodel/3d | unmodel/tripo3d, unmodel/fal |
| 🎵 Music generation | unmodel/music | unmodel/elevenlabs, unmodel/fal, unmodel/stability |
| 🔊 Sound effects | unmodel/sfx | unmodel/elevenlabs, unmodel/fal |
| 🔁 Voice conversion | unmodel/sts | unmodel/elevenlabs, unmodel/hume |
| 🎙️ Voice cloning | unmodel/voice-clone | unmodel/elevenlabs, unmodel/cartesia, unmodel/minimax |
| 🧪 Voice design | unmodel/voice-design | unmodel/elevenlabs, unmodel/fish-audio, unmodel/minimax |
| 🔌 Realtime audio config | none | unmodel/openai, unmodel/deepgram, unmodel/elevenlabs, etc. |
import { image } from "unmodel/image";
const request = image({
model: "openai/gpt-image-2",
prompt: "a lighthouse in fog",
aspectRatio: "16:9",
resolution: "1k",
});
JSON.stringify(request);
// → {"model":"gpt-image-2","prompt":"a lighthouse in fog","size":"1360x768"}
image({ model: "openai/dall-e-3", prompt: "...", quality: "hd" }); // ✅
image({ model: "openai/gpt-image-2", prompt: "...", quality: "hd" }); // ❌ TypeScript errorThe four newest surfaces are one line each: a clip, a still, a frame you want bigger, and an object.
import { lipsync } from "unmodel/lipsync";
import { avatar } from "unmodel/avatar";
import { upscale } from "unmodel/upscale";
import { threeD } from "unmodel/3d";
JSON.stringify(lipsync({ model: "veed/lipsync-2.0", source: { url: clip }, audio: { url: vo } }));
// → {"video_url":"https://ex.com/take.mp4","audio_url":"https://ex.com/vo.wav"}
JSON.stringify(avatar({ model: "fal/fal-ai/sync-lipsync/v3/image-to-video", image: { url: still }, audio: { url: vo } }));
// → {"image_url":"https://ex.com/face.png","audio_url":"https://ex.com/vo.wav"}
JSON.stringify(upscale({ model: "fal/fal-ai/clarity-upscaler", source: { url: still }, factor: 2 }));
// → {"image_url":"https://ex.com/face.png","upscale_factor":2}
JSON.stringify(threeD({ model: "tripo3d/v3.1-20260211", prompt: "a brass astrolabe", seed: 7 }));
// → {"model":"v3.1-20260211","prompt":"a brass astrolabe","model_seed":7}unmodel/3d is the first category that shipped with two providers, and on purpose: a 3D
vocabulary read off one vendor would be that vendor's schema with the names changed. The same
model through the aggregator compiles to a different body, which is the comparison it exists
to make cheap.
JSON.stringify(threeD({ model: "fal/tripo3d/h3.1/image-to-3d", image: { url: still } }));
// → {"image_url":"https://ex.com/face.png"}
JSON.stringify(threeD({ model: "tripo3d/v3.1-20260211", image: { url: still } }));
// → {"model":"v3.1-20260211","input":"https://ex.com/face.png"}unmodel/sfx is the newest, the smallest — four words — and the one where leaving a field out
is a decision. Omitting durationSeconds means the provider's own default, which is a
different number at every vendor and an HTTP 422 at one of them:
import { sfx } from "unmodel/sfx";
JSON.stringify(sfx({ model: "elevenlabs/eleven_text_to_sound_v2", prompt: "a door creaking open in a stone hall", durationSeconds: 4 }));
// → {"text":"a door creaking open in a stone hall","model_id":"eleven_text_to_sound_v2","duration_seconds":4}
sfx({ model: "fal/sonilo/v1.1/text-to-sound-effects", prompt: "rain on a tin roof" });
// ✅ compiles, and warns: approximated_param — this endpoint will generate 8 seconds
sfx({ model: "fal/cassetteai/sound-effects-generator", prompt: "rain on a tin roof" });
// ❌ TypeScript error: `durationSeconds` is required here — the wire answers 422 without itunmodel/sts is the newest, and the only category where most of the vocabulary is
required: a recording, a target voice, and the ref that picks the model. There is no prompt,
no length and no frame, because the answer to all three is "whatever the recording did".
import { sts } from "unmodel/sts";
const converted = sts({
model: "elevenlabs/eleven_multilingual_sts_v2",
audio: { file: recording },
voice: "21m00Tcm4TlvDq8ikWAM",
});
converted.request.url;
// → "https://api.elevenlabs.io/v1/speech-to-speech/21m00Tcm4TlvDq8ikWAM"
sts({ model: "hume/voice-conversion", audio: { file: recording }, voice: { name: "Male English Actor" } });
// → { audio: <Blob>, voice: { name: "Male English Actor" } }
// the same two words, and at Hume the voice is a form part rather than a URL segment
sts({ model: "elevenlabs/eleven_multilingual_sts_v2", audio: { file: recording } });
// ❌ TypeScript error: `voice` is required — a conversion with no target is not a conversionBecause audio is a required Blob at both providers, this whole category is library-only:
unmodel validate elevenlabs.sts tells you so rather than failing with "expected Blob".
Same pattern for every surface — inputs, formats, and extras narrow to the selected model. Per-category guides, including audio input routing, multipart helpers, and voice cloning: docs/surfaces.md. Full roster: docs/providers.md; per-provider TTS quirks: docs/tts.md.
✅ Validation
Invalid params throw UnmodelValidationError. Check it with UnmodelValidationError.isInstance(error) rather than instanceof — it survives a second copy of the package and a Worker realm boundary — and never retry it: validation is a pure function of the params, so the second attempt fails on the same issue as the first. Use .safe() to get issues back as values:
const result = chat.safe({
model: "openai/gpt-5.2",
messages: [{ role: "user", content: "Hello!" }],
});
if (result.ok) {
result.params; // validated provider body
result.warnings; // non-fatal validation findings
result.estimate; // input tokens and worst-case cost, when known
} else {
result.errors;
}What is checked:
- Shape, unknown fields, enums, and mutually exclusive params
- Model existence, deprecation, capabilities, and per-model exceptions
- Context, input, output, media, and provider-specific limits
- Estimated budget via
maxCostUSD - Unsupported or lossy unified translations
maxCostUSD turns the estimate into a gate — over budget is an error, not a warning:
const result = tts.safe({ model: "elevenlabs/eleven_multilingual_v2", text, voice }, { maxCostUSD: 0.01 });
result.ok && result.estimate.costUSD; // 0.0024.safeUnknown(), severity options, cost arithmetic, future model IDs, and the error/retry patterns for durable runtimes: docs/validation.md.
🔁 Retarget chat with .toApi()
Move one validated request to another provider that serves the same model:
import { chat } from "unmodel/anthropic";
const request = chat({
model: "claude-opus-5",
max_tokens: 4096,
thinking: { type: "enabled", budget_tokens: 2048 },
messages: [{ role: "user", content: "Explain retargeting." }],
});
const moved = request.toApi("openrouter");
moved.model; // "anthropic/claude-opus-5"
moved.request.url; // https://openrouter.ai/api/v1/chat/completions
moved.warnings; // id respelling or lossy translations
request.toApi("openai");
// ~~~~~~~~ TypeScript error: OpenAI does not serve ClaudeAuth moves with the provider; CHAT_AUTH from unmodel/chat maps each to its header name and scheme. .toApiSafe() is the non-throwing form; details and caveats in docs/validation.md.
🔁 Retarget media with .toApi("fal")
The same move for image, video and speech: fal re-serves other vendors' media models, so a validated native request can be sent to fal's queue instead.
import { video } from "unmodel/kling";
const request = video({
model_name: "kling-v2-5-turbo",
prompt: "A slow push-in through a rainy neon alley",
mode: "pro",
duration: "10",
});
const onFal = request.toApi("fal");
onFal.request.url; // https://queue.fal.run/fal-ai/kling-video/v2.5-turbo/pro/text-to-video
{ ...onFal }; // { prompt: "A slow push-in…", negative_prompt: "", duration: "10" }
onFal.warnings; // [] ← empty means the mapping was exact
video({ model_name: "kling-v1", prompt: "…" }).toApi("fal");
// ~~~~~ TypeScript error: fal serves no Kling v1Auth changes with the host — Kling takes authorization: Bearer <key>, fal takes authorization: Key <FAL_KEY>. A parameter fal cannot express is an error naming it, never a silent drop; a derived or snapped value is one warning. Six families across image, video and tts, with the mappings and the deliberate refusals in docs/providers.md.
🧾 Check responses
Provider check* helpers inspect raw responses and normalize quality and usage signals:
import { checkChat } from "unmodel/openai";
const payload = await response.json();
if (!response.ok) throw new Error(`Provider error ${response.status}`);
const report = checkChat(payload);
report.warnings; // truncation, filtering, refusals
report.finishReason; // normalized finish reason
report.usage; // input, output, cached, reasoning
report.costUSD; // actual cost from catalog ratesUse the checker from the provider that returned the response; media checkers follow their response documents (checkImages, checkTranscription, checkTts, …) — see docs/validation.md.
🔡 Types and values
import type { ChatParams } from "unmodel/types"; // emits no JavaScript at all
import { TTS_MODEL_PARAMS } from "unmodel/openai/values"; // runtime arrays for your <select>, ~1 KiB
TTS_MODEL_PARAMS["gpt-4o-mini-tts"].voices; // the exact array the validator enforcesTypes-only subpaths ship zero runtime, tested against a real build. Values entries re-export the same objects the adapter compiles with — asserted by ===, so a picker and the request it builds cannot disagree — and every import's bundle cost is measured and budgeted by a test. Full reference: docs/types-and-values.md.
📡 Send with anything
import OpenAI from "openai";
import { chat } from "unmodel/openai";
const request = chat({ model: "gpt-5.2", messages: [{ role: "user", content: "Hello!" }] });
await new OpenAI().chat.completions.create(request.toSdk("openai"));toRequestInit(request)for JSON + fetch; multipart results redirect you to theirtoFormDatahelper at compile time.toSdk("openai" | "google" | "google-vertex" | "ai-sdk" | …)for typed SDK handoff- Vercel AI SDK via
.toSdk("ai-sdk"), tools viawithJsonSchemaToolsfromunmodel/ai-sdk - WebSocket validators return a ready
wss://URL or a validated config object
Details: docs/integrations.md.
🗂️ Catalog and CLI
import { getModel } from "unmodel/catalog";
const model = getModel("openai", "gpt-5.2");
model?.limit.context;
model?.cost?.input;npx unmodel models openai gpt-5.2
npx unmodel validate openai.chat request.json
npx unmodel validate unified.image image.json --max-cost 0.05
npx unmodel validate unified.stt transcription.json --jsonvalidate exits non-zero for invalid params. Speech, image, and video catalogs are hand-maintained per provider (models from unmodel/elevenlabs, …) — see docs/integrations.md.
🏢 Providers
Every implemented provider has its own subpath with native field names, model IDs, routes, pricing, and quirks. Providers whose URL depends on your account expose factories (createAzure, createGoogleVertex, createAmazonBedrock, createCloudflare), and createOpenAICompatible covers proxies and self-hosted Chat Completions endpoints. Full roster and roadmap: docs/providers.md.
fal.ai
unmodel/fal covers 172 curated endpoints across ten verbs: image, imageEdit, video, lipsync, upscale, avatar, threeD, tts, stt, music. Four things here work unlike every other provider.
import { image } from "unmodel/fal";
const request = image({ endpoint: "fal-ai/flux/dev", prompt: "a cat", image_size: "landscape_4_3" });
JSON.stringify(request);
// → {"prompt":"a cat","image_size":"landscape_4_3"}
request.request.url; // "https://queue.fal.run/fal-ai/flux/dev"- The model is the route. The endpoint id is the URL path, so the selector is a pseudo-param named
endpoint, stripped before the body goes out. It cannot bemodel, becausemodelis a real wire field on some fal endpoints. Unified refs are unaffected:"fal/fal-ai/flux/dev"splits on the first slash. - Every request is a queue submit.
POST https://queue.fal.run/{endpoint}answers an envelope (request_id,status, and theresponse_url/status_url/cancel_urlto follow), not a file. Follow theresponse_urlfal hands back, never one you build. Polling stays with your transport code. - Auth is
Authorization: Key ${FAL_KEY}. TheKeyprefix is real and fal's own OpenAPI omits it, so unmodel states it in prose rather than deriving it. No unmodel export takes your key. - The types are generated from fal's own published OpenAPI.
bun run codegen:falrebuildssrc/providers/fal/gen/from committed per-endpoint snapshots. Curation, pricing and overlays stay hand-maintained indata/fal/, each row carrying a source URL, a date and a quote.
📚 Docs
- 📖 Surfaces guide — all fourteen categories, examples, quirks, bundles and custom packs
- ✅ Validation, cost, and response checks
- 🔡 Types and values reference
- 🔌 Integrations: fetch, SDKs, catalog, CLI
- 🗺️ Provider roster and roadmap
- 🗣️ TTS integrator's matrix
- 🏛️ Architecture decisions
🛠️ Development
bun install
bun test
bun run check
bun run build
bun run lint:pkg
bun run codegen # regenerate from the checked-in catalog snapshot
bun run codegen:refresh # refresh models.dev, then regenerateLicense
MIT
