@broberg/ai-sdk
v0.56.1
Published
Unified AI/LLM SDK — one facade, all providers, all capabilities, first-class cost control on every call.
Readme
@broberg/ai-sdk
One AI/LLM SDK — one facade, all providers, all capabilities, with first-class cost control on every call.
A provider-agnostic facade: your code calls ai.chat(), ai.vision(),
ai.image() — never a provider SDK directly. Swap providers by changing a tier,
not your call-sites. Every call returns a Usage (tokens, cost, latency,
transport, data-residency region) and can fan that out to any cost sink.
bun add @broberg/ai-sdk # or: npm i @broberg/ai-sdkWhere did the data go? — usage.region
Every response carries the data residency of the endpoint that actually answered:
const { text, usage } = await ai.chat({ prompt, tier: "smart" });
if (usage.region !== "eu") throw new Error(`personal data would have left the EU (${usage.region})`);Only "eu" is a positive claim. "unknown" means we cannot say, so
region !== "us" is not an EU check — it passes every single OpenRouter call.
region is derived from the host the request used, not from the provider's name —
vertex, azure and bfl are EU by default but each takes an override, so a
name-based table would report eu for a call someone had pointed at us-central1.
Three things that decide whether your guard works:
| | |
|---|---|
| After the call | usage.region — "eu" | "us" | "cn" | "unknown" |
| Before the call | regionOfHost(urlOrHostname) — the only way to refuse before bytes leave |
| Never | regionOfProvider(name) — see below |
import { regionOfHost } from "@broberg/ai-sdk";
if (regionOfHost("https://api.mistral.ai/v1") !== "eu") throw new Error("not EU");Do not build a residency guard on regionOfProvider. A name cannot answer for
hosting when the provider takes a baseUrl, so regionOfProvider("mistral") is
"unknown" — and a guard written as regionOfProvider(p) === "eu" therefore rejects
Mistral, the only EU route there is. Fail-closed and useless: it filters out the
thing it exists to allow.
An allowlist of (provider, model) pairs has the same hole from the other side: a
gateway in front of Mistral is still "mistral"/"mistral-large-latest", matches both
fields, and passes. regionOfHost answers "unknown" for it, which is the honest answer.
Quick start
import { createAI } from "@broberg/ai-sdk";
const ai = createAI(); // real adapters, keys from env (ANTHROPIC_API_KEY, …)
const { text, usage } = await ai.chat({ prompt: "Say hi in Danish" });
console.log(text, usage.costUsd);
const v = await ai.vision({ image: "https://…/photo.png", prompt: "Describe" });
const img = await ai.image({ prompt: "a sunlit beach in Blokhus" });
const da = await ai.translate({ text: "hello", to: "Danish" });
const emb = await ai.embedding({ text: ["a", "b"] });Capabilities
chat · vision · translate · image (fal.ai default / OpenRouter) · embedding · transcribe
(Whisper), plus prompt contracts with structured output:
import { z } from "zod";
const { data } = await ai.contracts.extract({
text: "Sanne is 40 and lives in Blokhus",
schema: z.object({ name: z.string(), age: z.number(), city: z.string() }),
});
// also: ai.contracts.{ mockup, design, classify, rerank }ai.judge — Jev (TypeSafe), F066. Not a chat model: typed yes/no, choice and
score questions about a piece of content, each answer with a calibrated probability
and a confidence you can gate on. $0.042 per million input tokens, output free.
US-hosted — not for personal data. See docs/API.md §4.
Read-aloud: what was actually SPOKEN is not your text
ai.tts({ pronunciations }) rewrites your text before the provider ever sees it —
broberg.ai becomes four spoken words. So text is a retelling of the audio, and a
word-highlighter built on text alone will drift at every rewritten word.
Two fields close that gap:
const { audio, wordTimings, ssml } = await ai.tts({
text, voice, wordTimings: true, pronunciations,
override: { provider: "azure" },
});
wordTimings.words // [{ text, startMs, endMs, sourceStart, sourceEnd }]
wordTimings.unaligned // spoken words we could NOT place — never a guessed offset
ssml // the EXACT markup we sent, verbatim — your ground truthsourceStart/sourceEnd index your original text, because Azure reports audio time
only and carries no text offset at all; the link back to the manuscript is derived here.
ssml exists so you can check that derivation instead of trusting it — it is the string
we sent, not a rebuild (filed by a consumer who could not verify what was spoken, F055.4).
It is undefined on routes that build no markup; an empty string would be a different claim.
Two matcher rules that surprise people, both measured in pronunciations:
- Case-insensitive.
ai,AiandAIall match a rule forAI. - A word touching a hyphen keeps its spelling.
broberg.ai-drevetis NOT rewritten unless that entry setsmatchInCompounds: true— the rule that keepsmailout ofe-mail. A consumer deriving their own expected-word list got exactly these two occurrences wrong before the field existed.
Word timings need Azure's BATCH route, a different call shape (submit → poll → ZIP),
and that route requires the resource's custom subdomain — the regional host answers 401
with a valid key, blaming the key. Set AZURE_SPEECH_RESOURCE; without it wordTimings
fails before the call, naming the fix.
Providers & tiers
Adapters: Anthropic (HTTP + claude -p subprocess), OpenAI, Google
Gemini, DeepInfra, OpenRouter (incl. MiniMax), fal.ai (images).
Calls route through named tiers — fast · smart · powerful · cheap · vision ·
embedding — each resolving to a (provider, model, transport) triple,
overridable per call:
await ai.chat({ prompt: "…", tier: "powerful" });
await ai.chat({ prompt: "…", override: { provider: "openrouter", model: "minimax/minimax-m2.7" } });Images — raster vs. vector.
ai.image()defaults to fal.ai (raster PNG). For vector/SVG output (logos), override to OpenRouter Recraft — the slug is an OpenRouter model, not a fal app-id:await ai.image({ prompt: "…", override: { provider: "openrouter", model: "recraft/recraft-v4.1-vector" } }); // → data:image/svg+xml;base64,… with ground-truth cost
recraft/recraft-v4.1-pro-vectorgives higher detail for complex vectors (~$0.30 per image vs ~$0.08);recraft/recraft-v4.1-prois the raster equivalent (~$0.21).
cheap defaults to the cheapest-that's-good-enough cloud model — Mistral Small
(EU/Paris-hosted, GDPR-safe, ~$0.10/$0.30) — so a cost-tier call is safe for
personal data by default; override per call for an even cheaper non-personal route.
(The claude -p subprocess transport is still available via explicit
override: { transport: "subprocess" }, but is no longer a default route.)
Where WOULD this go? — wouldRouteTo (a forecast, not a fact)
usage.region answers where a call went. It cannot answer may I send this?, because
it only exists once the bytes have left. wouldRouteTo answers that half:
import { wouldRouteTo } from "@broberg/ai-sdk";
const f = wouldRouteTo("smart");
// { kind: "forecast", provider: "mistral", host: "https://api.mistral.ai/v1",
// wouldRouteTo: "eu", onlyIf: [ …the assumptions… ] }It is deliberately not shaped like usage.region. There is no field called region on
it, and that is enforced by a test: a forecast that can be wrong must not wear the clothes
of a fact that cannot. onlyIf comes back with the answer rather than living in these
docs — the answer holds only while you override no baseUrl and pass no fallback, because
a fallback IS a route and the route decides residency.
"depends-on-config" is an answer, not a failure. azure, vertex, deepl, requesty
and fal build their host from values you supply. requesty is why this state exists: it
has both an EU and a non-EU host, so any region we named there would be a residency claim
decided by a table instead of by a route.
It refuses nothing. It never throws and never blocks a call — not even for a nonsense provider. Residency policy belongs to each consumer, not to this package; this is the information that makes that decision possible, which is the opposite of a gate.
Why it can say "eu" where regionOfProvider says "unknown": it asks the host the
SDK would actually use, not the provider's name. regionOfProvider("mistral") is
"unknown" by design — that adapter takes a baseUrl. Filed by a consumer who had
hand-copied our tier table into their own repo to answer this; that copy can now be deleted
instead of maintained.
Cost, budget & sinks
import { createAI, upmetricsSink, sqliteSink, multiSink } from "@broberg/ai-sdk";
const ai = createAI({
budget: { perCallUsd: 0.05, rollingUsd: 5 }, // pre-flight guard (throws BudgetExceededError)
costSink: multiSink([
upmetricsSink({ baseUrl: "https://upmetrics.org", apiKey: process.env.UPMETRICS_API_KEY!, agentName: "my-app" }),
sqliteSink({ dbPath: "./ai-cost.db" }), // Bun only — throws on Node, see below
]),
});Sinks: upmetricsSink (canonical), discordSink, sqliteSink, multiSink,
noopSink. A sink that fails during a call never crashes that call.
sqliteSinkandgetCostSummaryare Bun-only (changed in v0.48.0, released 22 September 2026). They are backed bybun:sqlite, which Node cannot import at all. On Node they now throw where you construct them — deliberately, and this is a behaviour change: up to v0.47 you got a sink object that threw on everyrecord(), and the client swallows per-call sink errors by design, so cost tracking went silently dead. An empty cost dataset and a working one look identical in a report. On Node useupmetricsSink.sqliteBudgetStoreis Bun-only for the same reason; it has always failed loudly there, because a budget error is not swallowed.
Cost-tracking is on by default (v0.24+)
You don't have to wire a sink. If UPMETRICS_API_KEY is in the env, a bare
createAI() auto-attaches the upmetrics sink — this exists because most
call-sites passed no sink, leaving ~91% of Mistral spend invisible. No key in
the env → no sink, no crash (ship-dark).
createAI(); // key in env → tracked; no key → nothing happens
createAI({ costSink: mySink }); // explicit sink wins
createAI({ costSink: null }); // explicit OPT-OUT — see below| Env var | Effect |
|---|---|
| UPMETRICS_API_KEY | The only switch. Present → tracking on. |
| UPMETRICS_AGENT_NAME | Row label. Set it — the fallback is npm_package_name, which is empty for a process not started via an npm/bun script (a node dist/… or Docker service logs as unknown). |
| UPMETRICS_BASE_URL | Ingest host. Defaults to https://upmetrics.org. |
| UPMETRICS_COMPLIANCE | 1 sets compliance mode. Usage carries no prompt/response text either way. |
Adopting it in a repo — two things to do first, or the adoption corrupts the numbers it was meant to reveal:
Opt out wherever you already report your own costs. Pass
costSink: null. Otherwise the SDK adds a second reporting path on top of yours and every call is counted twice, in production, with no error anywhere. (Found bybuddy, who aggregates tocli_usageand pushes hourly.)Keep the key out of your test run. Test runners auto-load
.env(Bun does; Vitest with a dotenv setup does), so any suite that reaches acreateAI()will POST fabricated usage into production telemetry. Note the trigger is importing a module that builds the client — a module-levelcreateAI()arms it without anything calling it. Strip the whole prefix:// test-setup.ts — bunfig.toml: [test] preload = ["./test-setup.ts"] for (const k of Object.keys(process.env)) { if (k.startsWith("UPMETRICS_")) delete process.env[k]; }Strip the prefix, not a list of names — a guard you must remember to update is one that silently rots. And assert the effect (no
UPMETRICS_*visible to a test) with a control proving the probe can go red; a test that checks "the key is undefined" passes while protecting nothing on a machine that never had the key.
Adoption is per repo, opt-in — there is no fleet-wide push. Turn it on when a repo has spend worth watching, with both guards in place.
License
FSL-1.1-Apache-2.0
