@bivelio/savings-layer
v0.5.2
Published
The savings layer for your LLM bill: a library inside your code that lets through only what has to be paid for and returns a SavingsReport per call, measured with your provider's own meter and kept apart from estimates. Commercial license, paid per seat,
Maintainers
Readme
In plain words
Every request your app sends to an AI is paid for by the text going in and out. Savings sits on that path and lets through only what's needed — skipping calls that don't have to happen, and trimming what does — so you pay for less. It is not a promise you take on faith: every saving we report is measured with your provider's own usage counter, the one your bill is built from (for a repeat served without calling the provider, the counter of the identical call that ran), and every saving — measured or estimated, and labelled as such — is handed back to you request by request.
Measure for free, then pay per seat — with no commission on your savings. The free Measure license (no card) runs the layer in shadow mode on your own traffic and shows what it would have saved, valued with your provider's counter. To turn the levers on, Pro is €15 / seat / month or €150 / seat / year, and Team is €129 / month with 10 seats included (€12 per extra seat; €1,290 / year, €120 per extra seat), EUR, VAT not included — a fixed fee, paid whether or not you save, and no commission on what you save. Enterprise (more seats, your own contract or invoicing terms) is by contract. And your prompts and completions never reach BiVelio: what BiVelio receives is per‑request usage — token counts, cost, the model id and opaque ids — never your content (LICENSE §5 lists every field; savings.bivelio.com/privacy explains each one and how long it is kept).
For engineers, in one line
@bivelio/savings-layer (BV‑SALA) is one layer you add inside your code, right
where your app calls the OpenAI‑compatible or Anthropic API. For every request it
picks the cheapest execution that still meets your quality bar and hands you an
auditable SavingsReport — the same numbers your dashboard reads. Traffic still
goes straight from your app to your provider; the SDK is never a middleman server.
Commercial software — © 2026 BiVelio Inc. Free to install, paid to save: BV‑SALA is fail‑closed and executes only with an active BiVelio license key. The free Measure license (no card) runs it in shadow mode: nothing you send is changed, nothing is served from cache, and you see what it would have saved. The license signature is verified offline; with
licenseServerUrlset, the SDK also callssavings.bivelio.comfor the revocation list, license renewal and usage metering. Your prompts, completions and provider keys never reach BiVelio — what travels is per‑request usage: token counts, cost, the model id, a request id and opaque seat and session ids (the exact list is in LICENSE §5). Use is governed by the LICENSE, which also sets out the fees, and the BiVelio Terms. Not a wrapper around any third‑party library.
Where it installs
Your application → BV‑SALA → provider API. Traffic always goes straight from your app to your provider — BV‑SALA is a library inside your code, never a middleman server, and it works with any OpenAI‑compatible endpoint (OpenAI, Azure, vLLM, SGLang, a gateway) as well as Anthropic.
Is this for me? If your product or automation calls the AI with its own API key — and pays a per‑token bill for it — yes: this is exactly where the saving installs. Only use ChatGPT or Claude? Then there is nothing to install: those products charge a flat subscription, so there is no per‑token bill to cut. The whole story, in plain words: Where does the saving install?
A coding agent with an API key? The package also ships a local measuring
gateway: bvsala claude (or bvsala codex, goose, qwen, opencode, kilo,
copilot with your own Anthropic key, aider) launches the agent through it and your real
spend shows up in your dashboard, token by token — nothing to configure, nothing
persisted. On exit it tells you how many requests it metered and which ones it
could not, and why. For Crush, Droid, Cline, Continue or Copilot in VS Code,
bvsala doctor prints the exact config block. Gemini CLI and Cursor cannot be
metered today, and the doctor says so. Like the SDK, the gateway is
fail‑closed: it starts only with your service key (BIVELIO_SERVICE_KEY) and an
active BiVelio license, which it collects with that key and verifies against the
key embedded in the package. Once running it never interrupts a request in flight:
if the license lapses, only new requests are refused (HTTP 402) until it is
renewed. It measures every request. With a PRO license it also applies one saving
by default: it shortens the cache-write TTL while your session stays hot, which
changes the price, never the content the model reads. The other levers (sibling
model, per-turn effort) stay off until you switch them on. On a Claude or ChatGPT
subscription there is no per-token invoice, so that traffic is metered at
list-price equivalent and never carries commission; the per-seat license applies
as for any other use.
The gateway forwards what your terminal sends. With no lever armed, your
provider receives the request body byte for byte and every header your client
sends except x-bvsala-report, which asks the gateway itself for a report (see
the limits): cache_control and its 1-hour TTL, tool search and
tool_reference blocks, the advisor, every anthropic-beta value. An armed lever
changes only what it declares, and a request it cannot rewrite exactly travels
untouched. A test suite checks this against a fake upstream on every change
(test/gateway-no-empeora.test.ts; the repository is private, so the linked
section lists exactly what it checks). The gateway also logs, once per session,
requests that come out worse than your client asked for: cache markers the
provider ignored, a 1-hour TTL written at 5 minutes upstream, and tool search off
(Claude Code turns it off by default behind any proxy, this one included;
bvsala claude turns it back on). What it checks, and its limits:
docs/terminales/README.md.
Python and any other language: the local gateway
The same gateway sits in front of any SDK that lets you set a base URL, whatever the language:
bvsala gateway --upstream openai --port 8403 # one gateway per upstream
bvsala gateway # Anthropic, the default, on 8402Flags and environment variables are in English (--port, --project, --force,
BVSALA_GATEWAY_EFFORT, BVSALA_GATEWAY_SIBLING_MODEL). The Spanish spellings from
0.4.x (--puerto, --proyecto, --forzar, BVSALA_GATEWAY_ESFUERZO,
BVSALA_GATEWAY_MODELO_HERMANO) were removed in 0.5: using one is an error that
names the English spelling, and the gateway does not start. bvsala --help lists
them all.
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8403/v1") # WITH /v1
stream = client.chat.completions.create(
model="gpt-4o-mini", messages=[{"role": "user", "content": "Hi"}], stream=True,
stream_options={"include_usage": True}, # without it a stream carries no usage and is not metered
)
for chunk in stream:
if chunk.choices: # the last chunk carries only the usage
print(chunk.choices[0].delta.content or "", end="")
from anthropic import Anthropic
claude = Anthropic(base_url="http://127.0.0.1:8402") # NO /v1With BIVELIO_SERVICE_KEY set, the gateway measures every request whose usage the
provider reports — Anthropic Messages, OpenAI Chat Completions and Responses, streamed
or not — for models with a list price in its catalog, priced at that list price. A
request it cannot meter (a model without a list price, a body it cannot read, a stream
that never asked for usage) is still forwarded, and its log line says so and why. That
is measurement, not savings: the saving levers of the SDK are TypeScript, and the
gateway's own levers apply to Anthropic /v1/messages traffic with a PRO license.
bvsala doctor prints the recipe for each SDK.
Each request can also bring its report back to your code. Send the header
x-bvsala-report: 1 (in Python, default_headers={"x-bvsala-report": "1"} on
the client) and the gateway answers with x-bvsala-request-id. On the same
gateway, GET /_bvsala/v1/reports/<that id> returns what it recorded for your
dashboard: metered with the SavingsReport (the provider's token counts, the
cost at list price and cost.basis), or not_metered with the reason from the
log line. ?response_id=<the provider's response id> finds it too, and
wait_ms waits for a request still in flight. The header never reaches your
provider, and a request without it is not changed at all:
the details.
Check it before you install
What you can check before paying is what we have measured: the published benchmark, trial by trial, with the raw CSV linked, at savings.bivelio.com/en/how-it-works, and the free Measure license, which runs the layer in shadow mode on your own traffic and shows what it would have saved, valued with your provider's counter. The verifier that replays the optimizer on your own data lives in your dashboard once you sign in.
Upgrading from 0.4? Every breaking change of 0.5, with a before and after for each, is in Migrating from 0.4 to 0.5.
Set it up in one command
npx @bivelio/savings-layer@latest initIt reads your package.json (Next.js, Express, the Vercel AI SDK, Mastra,
LangChain.js, the Anthropic SDK or the OpenAI SDK) and prints the three steps below,
written for your stack — the install line for your package manager, which keys are
missing, the exact code for your adapter — plus a free check: a question the layer
answers on your machine, so the model is not called. It reads package.json files only
(yours, and those in node_modules for the installed versions), never a .env file or a
key's value; it uses no network of its own and writes nothing.
--write creates the new files after asking you, and never overwrites one; --json
prints the same plan for a script or a coding agent.
Install in three steps
1 · The package
pnpm add @bivelio/savings-layer gpt-tokenizer
# or: npm i @bivelio/savings-layer gpt-tokenizer · yarn add @bivelio/savings-layer gpt-tokenizergpt-tokenizer is optional, but install it if you call OpenAI models: it is the
provider's own tokenizer, and the SDK loads it by itself. Without it, counts fall back
to a built‑in estimate and the reconstructed savings are less accurate. Either way, a
saving where the layer removed something is reported as estimated, and estimated
savings are never billed — the SDK warns once when the tokenizer is missing.
2 · The license — from your dashboard (free Measure license, no card; Pro €15 / seat / month or €150 / seat / year; Team €129 / month with 10 seats included — EUR)
export BIVELIO_LICENSE_KEY=… # your license key
export BIVELIO_SERVICE_KEY=… # service-account key: auto-renewal + seat metering (required by the gateway)3 · Wrap the call — this is the entire code change. The package is ESM only:
put this in a "type": "module" project or a .mts file.
import { createSavingsLayer, OpenAICompatibleProvider } from "@bivelio/savings-layer";
// Your provider. LLM traffic goes straight to it — the SDK never proxies it.
// Any OpenAI-compatible endpoint works: OpenAI, Azure, vLLM, SGLang, a gateway.
// For Gemini, use the native `GoogleProvider` instead (see "Gemini: the native adapter").
// Model ids are qualified (`openai/gpt-4o-mini`); the prefix is stripped on the
// wire for OpenAI, Azure and Google's Gemini API and kept for gateways — see
// `stripProviderPrefix`.
const provider = new OpenAICompatibleProvider({
baseUrl: "https://api.openai.com/v1",
apiKey: process.env.OPENAI_API_KEY,
});
const layer = createSavingsLayer({
provider,
licenseKey: process.env.BIVELIO_LICENSE_KEY, // required — from your savings.bivelio.com dashboard
licenseServerUrl: "https://savings.bivelio.com", // revocation list, renewal + metering — omit it and the two lines below do nothing
serviceAccountKey: process.env.BIVELIO_SERVICE_KEY, // seat metering + AUTO-RENEWAL pickup
});
const { text, report } = await layer.generate({
model: "openai/gpt-4o-mini",
messages: [{ role: "user", content: "What is 2 + 2 * 3?" }],
});
console.log(text); // "8" (computed locally — no LLM call)
console.log(report.calls.avoidedLlm); // → 1: a whole call you never paid for
console.log(report.cost.netSavingsRatio); // → 0..1 ratio of the baseline you did not pay (1 here)
console.log(report.cost.netSavingsUsd); // → the same saving in USD (baselineEstimate − actual)
console.log(report.cost.basis); // → "estimated": nothing ran, so there is no provider counterThat's it. From the first request, report carries the saving and says how it
was obtained: actual is what the provider metered; basis is "measured" when
the request went through untouched, "reference" when a repeat of a byte‑identical
request was served without calling the provider — its baseline is the provider's own
counter for the identical request that already ran and that you already paid
(report.cache.reference points to it; with an output schema, only when your schema
traveled exactly as sent, in strict mode, in the provider's native response‑format
parameter, with nothing added to the prompt), valued at the provider's cached‑input
rate, the lowest it bills for repeating a prompt; the output counts only under deterministic
sampling (temperature: 0), on a model with a list price in the catalog that does not
reason and honors that setting, and when the call that ran carried no hidden reasoning (the
provider reported zero reasoning tokens or, where it does not break them out, its answer
had no reasoning block), and then at 95 % of its counter, a margin for the small drift
temperature: 0 still has —
"estimated" when the layer removed something (the baseline is then reconstructed
with the model's tokenizer, or with a calibrated estimate where no exact tokenizer
exists, as for Anthropic models), and "unmeasured" when the provider returned no
usage. Only measured and reference figures reach your invoice; estimated ones
never do.
A repeat is valued by reference only when the call that ran reached the provider's
own API (https://api.openai.com, https://api.anthropic.com or
https://generativelanguage.googleapis.com), because the list price it is valued at
is the invoice only there. When OpenAICompatibleProvider, AnthropicProvider or
GoogleProvider points anywhere else — a proxy such as LiteLLM, a gateway, Azure, Bedrock, Vertex —
or when the Anthropic SDK or LangChain adapter reads that its client does, that call
leaves no reference: its repeats stay estimated, and the reference trace of the
call that ran says unofficial-origin. When the layer cannot tell where the call
went (a provider of your own, the AI SDK and Mastra adapters, a LangChain model that
does not expose its base URL), the rule applies as it always has. The layer reads the
configured base URL, not the wire: a custom fetch that sends the request elsewhere
is not seen.
Provider refusals. A refusal is measured like any other call and is never cached,
repaired or replayed by reference. One kind is priced at 0: an Anthropic
stop_reason: "refusal" that arrives before any output (no content block, zero output
tokens) with stop_details.category set to cyber, general_harms or null. Anthropic
documents these as
not billed.
Its tokens still appear in report.tokens. cost.actual, cost.baselineEstimate and
cost.netSavingsRatio are all 0, so the request credits no saving: not measured, not
estimated and not by reference. A concurrent identical request served from it credits
nothing either. The same rule applies in the gateway. Any other refusal is priced from
its usage as usual. That covers bio, frontier_llm and reasoning_extraction, a
refusal after partial output, and a category the table does not list.
Most refusals carry no signal from the API: the model writes "I can't help with that."
as ordinary text and ends with end_turn or stop. The layer also looks for these,
with a conservative text heuristic. It checks only the opening of the answer, or the
whole answer when it is short, and it needs a real refusal clause, such as "I'm sorry,
but I can't help with that", "I cannot assist", "I can't determine…" or "As an AI, I
cannot provide". After an apology, any verb counts ("I'm sorry, but I can't
determine…"), and so does a stated reason with no "I can't" ("I'm sorry, but
accessing … is illegal", "… would be a violation of…"). When the answer opens with
the refusal itself ("I can't", "I won't", "I'm not going to", "I'm unable to", and
"No puedo", "Je ne peux pas", "Ich kann … nicht"…), any verb counts too ("I can't
argue that…"), except for a tested list of idioms that are not refusals ("I can't
wait", "I won't lie", "I can't find any issues", "I can't be sure, but…").
A single word is not enough.
An answer that opens with empathy and, in its first two sentences, sends the user to
a mental health professional, a crisis line, emergency services or someone they
trust, instead of answering, counts too (rule redirection). It covers English,
Spanish, Catalan, French, German, Italian and Portuguese. A match is measured and
served as usual. It is not written to the exact or semantic cache, and it is not
replayed by reference. A concurrent identical request served from it is estimated.
The trace says why (skipReason: "text_refusal", with the rule). A refusal
already stored by an earlier version is not served either. In the gateway, the
coalesced twin of such a response is estimated (primary-text-refusal). A false
positive costs one cache hit. A normal answer that starts with "Sorry for the delay,
here is…", praises with "I can't recommend this enough", has "I can't" later on, or
offers empathy with practical help and no referral ("I'm sorry to hear about your
flight; here are three options…") is cached as before.
Sensitive topics. A request whose user turns touch self-harm or suicide, a mental
health crisis or a medical emergency ("How do I commit suicide?", "I'm having a panic
attack", "my friend is not breathing") is never read from or written to the exact or
semantic cache, even when the answer is useful, and leaves no reference record. It is
measured and served as usual; a concurrent identical request served from it is
estimated. The policy and cache-write traces say why
(skipReason: "sensitive_topic", with the category). The check is lexical and
conservative: phrases, not single words ("kill a Python process", "Suicide Squad" and
"cut my hair" do not count). It covers the same twelve languages as the freshness
and personalization checks, with shorter lists outside English, Spanish, Catalan,
French, German, Italian and Portuguese. In the gateway, the coalesced twin of such a
request is estimated (sensitive-topic). A false positive costs one cache hit.
Streaming. generateStream() takes the same request and hands you the answer as
the provider writes it, then the same SavingsReport:
const stream = layer.generateStream({
model: "openai/gpt-4o-mini",
messages: [{ role: "user", content: "Summarize the release notes in one line." }],
});
for await (const event of stream) {
if (event.type === "text") process.stdout.write(event.delta);
}
const { report: streamed } = await stream.result;
console.log(streamed.cost.basis); // same rules as generate(), read from the provider's final usage
console.log(streamed.streaming?.delivery); // "live", or "whole" when the answer arrived in one pieceThe SDK asks the provider for its final usage at the end of every stream (on
OpenAI‑compatible endpoints, stream_options.include_usage); a stream that ends
without it is reported as unmeasured and never billed. With an output contract, or
a request declared deferrable, the answer is produced exactly as generate() would
and arrives as a single text event. Breaking out of the loop while the provider is
still writing cancels its request, and nothing is cached or recorded; once its answer
is complete, or when it arrives in one piece, nothing is cancelled.
await stream.result tells you which happened.
Inside your framework
Already on an agent framework? The layer plugs in at that framework's own extension point, from a subpath of this same package. The four adapters behave alike:
- They answer without calling the model when they can — a computed answer (the arithmetic above), an exact repeat of a request that already ran, an identical request already in flight — and otherwise let your framework make exactly the call it was going to make. They never rewrite what your framework sends.
- Tools are treated as having side effects unless you list them in
readOnlyTools: a request that offers any other tool is never answered from a cache, because the action has to run. - Fail‑closed, like the SDK: without an active license the call fails with
LicenseError(a framework may wrap it, keeping its name and message) and the model is not called. Put the adapter last, so nothing sits between it and the model.
Every model call gets the same SavingsReport as with the SDK — its basis says
whether a figure comes from your provider's own counter — and onReport receives it
too.
Vercel AI SDK
pnpm add @bivelio/savings-layer ai @ai-sdk/openai — a LanguageModelMiddleware
for ai 6 or 7:
import { generateText, wrapLanguageModel } from "ai";
import { openai } from "@ai-sdk/openai";
import { savingsMiddleware, getSavingsReport } from "@bivelio/savings-layer/ai-sdk";
const savings = savingsMiddleware({
licenseKey: process.env.BIVELIO_LICENSE_KEY,
licenseServerUrl: "https://savings.bivelio.com",
serviceAccountKey: process.env.BIVELIO_SERVICE_KEY,
readOnlyTools: ["get_weather"],
});
const model = wrapLanguageModel({ model: openai("gpt-4o-mini"), middleware: savings }); // last in the list
const result = await generateText({
model,
prompt: "What is 2 + 2 * 3?",
providerOptions: { bivelio: { context: { sessionId: "user-42" } } },
});
console.log(getSavingsReport(result)?.cost.basis); // "estimated": answered without calling the modelFor streamText, read it with getSavingsReport({ providerMetadata: await result.providerMetadata }).
Mastra
pnpm add @bivelio/savings-layer @mastra/core — an input processor for
@mastra/core 1.x:
import { Agent } from "@mastra/core/agent";
import { savingsProcessor } from "@bivelio/savings-layer/mastra";
const bivelio = savingsProcessor({
licenseKey: process.env.BIVELIO_LICENSE_KEY,
licenseServerUrl: "https://savings.bivelio.com",
serviceAccountKey: process.env.BIVELIO_SERVICE_KEY,
onReport: (r) => console.log(r.cost.basis),
});
const agent = new Agent({
id: "support",
name: "Support",
instructions: "Answer briefly.",
model: "openai/gpt-4o-mini",
inputProcessors: [bivelio], // last in the list
});
console.log((await agent.generate("What is 2 + 2 * 3?")).text); // "8"It covers generate() and stream(), not generateLegacy()/streamLegacy() or a
durable agent that resumes in another process; identical requests in flight are not
shared here. emitReportPart: true also streams each report as a
data-bivelio-savings-report part.
LangChain.js
pnpm add @bivelio/savings-layer langchain @langchain/core @langchain/openai — an agent
middleware for createAgent in langchain 1.x. The model string ("openai:…") is loaded
through its integration package, so @langchain/openai (or @langchain/anthropic) has
to be installed:
import { createAgent } from "langchain";
import { savingsAgentMiddleware, getSavingsReportFromMessage } from "@bivelio/savings-layer/langchain";
const agentSavings = savingsAgentMiddleware({
licenseKey: process.env.BIVELIO_LICENSE_KEY,
licenseServerUrl: "https://savings.bivelio.com",
serviceAccountKey: process.env.BIVELIO_SERVICE_KEY,
});
const lcAgent = createAgent({ model: "openai:gpt-4o-mini", tools: [], middleware: [agentSavings] }); // last
const state = await lcAgent.invoke({ messages: [{ role: "user", content: "What is 2 + 2 * 3?" }] });
console.log(getSavingsReportFromMessage(state.messages.at(-1))?.cost.basis);The model call is measured when it finishes; tokens you stream with
streamMode: "messages" keep coming from the model as usual. A chat model with its
own cache set is reported as unmeasured: its counter may be a cached copy.
Anthropic TypeScript SDK
pnpm add @bivelio/savings-layer @anthropic-ai/sdk — a client middleware for
@anthropic-ai/sdk 0.103 or later:
import Anthropic from "@anthropic-ai/sdk";
import { anthropicSavings } from "@bivelio/savings-layer/anthropic-sdk";
const claudeSavings = anthropicSavings({
licenseKey: process.env.BIVELIO_LICENSE_KEY,
licenseServerUrl: "https://savings.bivelio.com",
serviceAccountKey: process.env.BIVELIO_SERVICE_KEY,
});
const client = new Anthropic({ middleware: [claudeSavings.middleware] }); // last in the list
const message = await client.messages.create({
model: "claude-opus-5",
max_tokens: 256,
messages: [{ role: "user", content: "What is 2 + 2 * 3?" }],
});
console.log(claudeSavings.reportFor(message)?.cost.basis);It runs once per HTTP attempt, inside the SDK's retries: an attempt the provider
rejects reaches the SDK unchanged and is not recorded. Only message creation goes
through the layer; countTokens, batches, models and files pass untouched.
The middleware forwards your request body unchanged, history included, so it never
edits a turn before a thinking block. It also does not remove parameters the model
rejects. On Claude Sonnet 5.5, Opus 5.5, Fable 5.1 and Mythos 5.1 a forced
tool_choice (any or tool) returns a 400. On those models, and on Opus 4.7/4.8,
Opus 5, Sonnet 5 and Fable 5, so does a non-default temperature, top_p or top_k.
The middleware warns once per model and cause (silenceWarnings turns it off) and
the API's 400 reaches you unchanged.
These models reject some request parameters with a 400 (forcing tool
use,
Sonnet 5.5).
AnthropicProvider sends a request they accept instead of failing:
- Output contract. A forced
tool_choicereturns a 400 on these models, with or without thinking. When your schema meetsstrictmode, the contract travels as structured output (output_config.format). Without your tools, the provider still enforces the contract. With your tools, they stay callable withtool_choice: autoand the final text answer comes back in the contract's shape. When your schema does not meetstrict, the contract travels as a tool the model is not forced to call, with one sentence in the system prompt that tells it to answer through that tool. Models that support forced tool use (Opus 5, Sonnet 5 and the 4.x models) get the same request as before. temperature. A non-default value is dropped, with one warning per model, because it can only return a 400. The answer is then sampled, not deterministic.temperature: 1, the API default, is sent as is.- A model the table does not list (a new model or a gateway alias): when the API
rejects either parameter, the provider repeats the request once without it and
remembers the model for later requests. A 400 is not billed, so the retry is not
counted in
retries. - Thinking blocks.
ChatMessagedoes not carry thinking blocks, so this adapter never sends one. The preserved-thinking check, which returns a 400 when a turn before a thinking block has been edited, has nothing to check on its requests, even when the levers that rewrite history are on.
What you get on every request
A SavingsReport with the numbers behind every claim, kept in four separate
ledgers that are never summed into one flattering figure:
| Ledger | Meaning |
| --- | --- |
| avoided_calls | LLM / retrieval / tool calls that never executed |
| eliminated_tokens | information dropped because it was not needed |
| compressed_tokens | information kept, but encoded with fewer tokens |
| provider_cached_tokens | tokens still in the prompt, billed/processed more cheaply |
Your provider's own cache discount, shown apart and never billed. With prompt
caching, most of what you save often comes from your provider's prefix cache
(Anthropic cache_control, OpenAI's automatic cache, Gemini's implicit cache), not
from the layer. cost prices cached tokens at the cached rate on both sides, so that
discount is never credited as the layer's saving. report.nativeSavings shows it on
its own, from your provider's counter at list price, so the report adds up to your
invoice:
report.nativeSavings;
// {
// provider: "anthropic", mechanism: "explicit", layerMarked: true,
// cacheReadTokens: 40000, cacheWriteTokens: 3000, cacheWrite1hTokens: 0,
// inputPer1k: 0.003, cachedInputPer1k: 0.0003, writeMultiplier: 1.25, write1hMultiplier: 2,
// readDiscountUsd: 0.108, // reads × (input − cached-input rate)
// writePremiumUsd: 0.00225, // writes × input × (multiplier − 1)
// netUsd: 0.10575, // can be negative: a cold write pays before any read
// basis: "measured",
// }It is present only when a call ran and the provider reported cache reads or writes;
a cache hit or a coalesced request never carries it. It is not in cost.netSavingsRatio,
not in cost.baselineEstimate, and never in the commission base. Your dashboard shows
it as its own card, Native savings unlocked, next to the credited savings.
Quality, kept apart from money. report.qualityLedger records how the answer you
received earned its place: whether it came from your provider, a replay, or the answer
to a reworded question (with its similarity score); whether it was checked against your
output contract, repaired or cut off by the output limit; and what the model cascade
decided. It holds no money, is never added to the four ledgers, and never leaves your
process. Reusing answers to reworded questions is opt-in, and by default it covers
single-shot requests only: requests that declare tools, carry tool calls or results, or
hold more than one user turn are left out unless you set
savings.semanticCache: "include-agentic".
Calibrated reuse (opt-in). A fixed similarity floor means something different for
every embedder and every workload, so you can make the cache earn its floor on your own
traffic instead: new InMemorySemanticCache({ embed, calibration: { maxErrorRate: 0.02 } }).
It then serves no reworded answer until your labels show, at 95 % confidence, that at
most 2 % of the answers it would reuse from some score up are wrong, and it re-tests that
floor as labels accumulate. Until then every candidate goes to your provider, exactly as
it would without the cache, and the fresh answer is compared with the stored one by your
judge (by default only identical answers count, which suits classification and
structured output; pass your own for free text). Each calibrated reuse carries its bound
and the labels behind it in report.qualityLedger.cacheMatch.calibration. Calibration
costs reuses, never an extra call, and it does not change how anything is billed. If you
already have labeled pairs, calibrateSemanticFloor() runs the same test offline.
And a promise you can hold it to: never worse than your JSON. Payload encodings
are chosen by counting real tokens with the destination model's own tokenizer, and
BV‑SALA only moves away from plain JSON when the saving clearly clears the bar —
report.serialization shows what was chosen and what each candidate would have cost.
The per-request figures above are what the layer can show without running your request twice. To see what it does on your own traffic, with both sides metered by your provider, turn on verification sampling:
const checkedLayer = createSavingsLayer({
provider,
licenseKey: process.env.BIVELIO_LICENSE_KEY,
licenseServerUrl: "https://savings.bivelio.com",
serviceAccountKey: process.env.BIVELIO_SERVICE_KEY,
verification: { sampleRate: 0.02 }, // about 2 in every 100 eligible requests
});For that share of eligible requests, after your answer has been returned, the SDK sends the same request once more to your provider exactly as it would go without the layer — the model you asked for, all your tools, your data inline, no cache marks — and discards the second answer. Each pair (what the unoptimized copy cost and what the optimized request cost, both from your provider's usage counters) appears under Proof › Your bench in your dashboard, with the mean difference, its 95 % interval and every pair as a CSV download.
- It costs you money. You pay your provider for every copy: roughly the sampling
rate times what that traffic would cost without the layer. The SDK says so once
per process when it starts (
silenceWarnings: truemutes that notice too). - It is never billed. Pairs travel on their own channel into their own table. They never enter measured savings, your statement or any commission.
- Only requests served intact are checked. Requests that took the fail-open path,
were cut off, missed their output contract or came back without a provider counter
are counted as excluded and never sent twice.
savings: { verification: false }excludes a single request. - Both sides follow the same rule. A pair is compared only when the copy also finished normally and met your contract; a copy cut off by the output limit is counted and exported, never scored as a saving.
- The headline is the floor, and it waits. It is the lower bound of the 95 % interval, shown once there are enough pairs overall and groups with too few pairs of their own are a small share of your traffic, so a group where the layer costs you money cannot be left out of it.
- Only counts leave your infrastructure. The copy goes to your own provider, like the original; BiVelio receives token counts, costs, model ids and outcomes — never prompts or answers.
close()settles every copy still running: it is recorded as abandoned, never as a pair. On serverless platforms, pass your runtime'swaitUntilasverification.waitUntilso checks can finish.
Two request flags ask your provider for a lower price on the same tokens. Neither
changes the answer, and what they are worth is reported next to the four ledgers:
never folded into cost.*, never part of any commission.
savings.deferrable asks for the provider's Flex tier: OpenAI, and Gemini through
GoogleProvider (serviceTier: "flex") or its OpenAI‑compatible endpoint. Anthropic has no cheaper synchronous tier: its
discount lives only in Message Batches, which layer.batches (below) sends to.
Bedrock and DeepSeek off‑peak pricing are not supported.
- Latency. OpenAI publishes no latency commitment for Flex; Google publishes a
target of minutes. Keep deferrable requests off user‑facing paths.
report.deferred.latencyquotes what the provider documents. - Risk. Without capacity the provider refuses the request (429, or 503 on Gemini)
instead of waiting; OpenAI documents that the refusal is not charged. With
priceLevers: { onCapacityError: "standard" }the layer repeats it once at the standard tier and price, and says so inreport.deferred.capacityFallback. - What it is worth. Your provider's usage counter for that call, priced at the two
published rate cards (standard, and the tier that actually served it), component
by component, with the date the prices were checked. It is
measuredonly when the served tier was read back from the provider's own API, the prices are list prices and the counter says how much of the input was read from (or written to) the prompt cache; otherwiseestimated, andbasisReasonsays why. It never exceeds the difference between the two published rate cards, even with your owncostModels. A tier the provider did not confirm is never credited.
savings.standalone (GPT‑5.6 and later, on OpenAI's own API) declares that nothing
after the stable prefix of this request will be sent again. The provider then writes
its prompt cache only at the end of that prefix, or nowhere, instead of at the end of
your message, which on these models costs more than plain input. The trade‑off:
repeats and continuations of this exact prompt lose the cache reads the provider
default would give them, so leave it off conversation turns you will continue. Its
figure, report.cacheWrites.avoidedWritePremium, is always estimated: the provider
default was not run.
const leveredLayer = createSavingsLayer({
provider,
licenseKey: process.env.BIVELIO_LICENSE_KEY,
priceLevers: { deferrable: true, onCapacityError: "standard" }, // layer-wide defaults
});
const { report } = await leveredLayer.generate({
model: "openai/gpt-5.6-luna",
messages: [
{ role: "system", content: "…your long, fixed instructions…" },
{ role: "user", content: "Classify ticket 42." },
],
savings: { standalone: true }, // a request's own value always wins
});
report.deferred?.tierSaving; // { amount, basis, standardCost, appliedCost, pricesCheckedOn, basisReason }
report.cacheWrites; // { mode, breakpoints, avoidedWritePremium?, reason }layer.batches sends requests that can wait hours to the provider's own batch API,
billed at its published batch rate card: Anthropic's Message Batches through
AnthropicProvider, and OpenAI's Batch API through OpenAICompatibleProvider on
api.openai.com. Gemini's batch mode is not supported yet.
- What runs. Each request is translated by the same provider adapter as
generate()and sent as you wrote it: the layer's own per‑request savings (caching, compression, model cascade, retrieval) do not run on batches. The layer splits the job at the provider's limits (and, on OpenAI, one model per file). - Latency. The provider's: results arrive within its batch window, not on a
request/response path.
submit()returns a plain‑JSON job you can store and read from another process. - What it is worth. Each result's usage counter at the standard and at the batch
published rate cards, component by component, with the same
measured/estimatedrule as above. The tier is read from the provider: Anthropic reports it on every result; on OpenAI it is the batch the result came from. Nothing is metered and the batch saving is never billed.
const tickets = [{ id: "t-1", text: "My invoice is wrong." }]; // your deferrable work
const job = await layer.batches.submit(
tickets.map((t) => ({
customId: t.id,
request: { model: "openai/gpt-4o-mini", messages: [{ role: "user", content: t.text }] },
})),
);
// …later, even from another process (store `job` as JSON):
const results = await layer.batches.results(job);
results.ended; // false while the provider is still working
results.items; // one per request: { customId, outcome, text, usage, tierSaving?, reason }
results.summary; // { succeeded, errored, expired, credited, tierSaving, basis, pricesCheckedOn, reason, … }In an agent loop your history is re-sent on every turn and your provider reads it from its prompt cache at a fraction of the input price. Rewriting anything it has already cached turns everything after it back into full-price cache writes, which is how "fewer tokens" ends up as a bigger bill. So this lever only ever changes a tool output on the turn it arrives, and re-sends exactly the same bytes afterwards.
It removes terminal colour codes, superseded progress-bar states, whitespace between JSON tokens and runs of identical lines (folded into one line plus a count), and nothing else: no summaries, no dropped words. Each change either reverses exactly or leaves what a terminal would display. Outputs of calls that name a file are left verbatim, so an agent that edits by exact match still sees the file as it is.
const agentLayer = createSavingsLayer({
provider,
licenseKey: process.env.BIVELIO_LICENSE_KEY,
toolOutputCompaction: true, // or { exceptTools: ["read_logs"], minSavedChars: 512 }
});
// or per request: savings: { toolOutputCompaction: true }- Switch it on for whole conversations. The decision depends only on each output's own text, so every process and every turn produces the same bytes. A conversation that started without it is rewritten once, on its first compacted request.
- Counted, not billed. What it removes shows in
report.tokens.compressedInputand inreport.traces; it never enterscost.*or any commission. Your provider's invoice shows the effect. - Inside a framework, call
compactToolOutput(text)in your tool before you return its result: your history then holds the compacted text itself. - Measure it by conversation, not by request. A verification copy re-sends one request with the original history, which your provider has not cached, so on turns that re-send compacted outputs the copy costs more than going without the layer would have. Compare whole conversations with and without it.
By default the layer writes the provider's prompt cache at two fixed places: the end
of the system header and the end of the history before your last user message, and
stops marking a prefix whose measured hit rate does not pay the write premium. On
Anthropic, where a write costs 1.25x, the history and a tool catalogue sent without a
system header are marked only when the request itself shows they will be sent again:
a tool-loop step after your last user message, or a conversation already past its
second turn. The first follow-up of a conversation ([question, answer, question])
marks only the system header, so two-turn chats pay no premium that nothing reads.
In an agent loop with no system header, the end of the tool list is marked once the
catalogue reaches the model's minimum cacheable size (512 tokens on Claude Sonnet 5.5
and Opus 5.5, 1,024 on Sonnet 5, 4,096 on Haiku 4.5). That discount is the provider's
cache on your bytes: it is never credited as the layer's saving.
cacheBreakpoints: "auto" places them from what your traffic actually does. The layer
compares each request's prefix with the ones it sent before, per model, tool set and
system header, and learns how often each part is sent again and after how long. That
costs nothing: it compares locally and marks nothing to find out. Once the evidence
is conclusive:
- Tail. When requests continue each other (agent tool loops, conversations), it also marks the last block, so the next turn reads the whole prompt from cache. Tool results arrive after your last user message, so the fixed placement never caches them.
- History. When the history before the question is not sent again (one-off requests with their own document), it moves the breakpoint back to the system header, so no write premium is paid on a prefix nobody reads.
- One-hour lifetime (Anthropic). When turns come back after more than five minutes but within the hour, and the size of the request makes the 2x write pay.
- Explicit mode (GPT‑5.6 and later, on OpenAI's own API). When continuations are
rare, your question is not written to the cache.
savings.standalonestill wins when you set it, andreport.cacheWritessays when the layer chose it.
Until then it behaves exactly like the default. Every decision and its reason is in
report.traces (prefix-cache-placement), and the effect is read from the provider's
own counters: if they do not confirm the cache reads the layer predicted, that
traffic goes back to the default placement. The write premium it causes, one-hour
writes at 2x included, is charged to cost.actual; the cache discount it earns is
counted in ledger 4 like any provider cache read, never credited as the layer's
saving and never part of any commission.
const placedLayer = createSavingsLayer({
provider, // Anthropic, or OpenAI GPT-5.6+ on OpenAI's own API
licenseKey: process.env.BIVELIO_LICENSE_KEY,
cacheBreakpoints: "auto",
});Calibrated reuse and cacheBreakpoints: "auto" learn from your traffic, and by default
what they learn lives in the process. On serverless platforms (Vercel Functions, AWS
Lambda) an instance may serve one request or live a few minutes, so neither would ever
get past its starting point. Give each one somewhere to keep it. Giving the layer a
semanticCacheStore only says where reworded questions are kept; semanticCacheDefault: true
is what switches reuse on for every request (a request's own savings.semanticCache
still wins), and without it, or savings.semanticCache on each request, the cache
stays empty and learns nothing:
import { InMemorySemanticCache } from "@bivelio/savings-layer";
// Any key-value client your instances share, such as Redis.
type SharedKeyValue = {
get(key: string): Promise<unknown>;
set(key: string, value: string, options?: { px: number }): Promise<unknown>;
};
async function serverlessLayer(kv: SharedKeyValue) {
const semanticCache = new InMemorySemanticCache({ calibration: { maxErrorRate: 0.02 } });
const saved = await kv.get("bvsala:semantic-calibration");
if (saved) semanticCache.importCalibration(saved); // the object or its JSON text
const layer = createSavingsLayer({
provider,
licenseKey: process.env.BIVELIO_LICENSE_KEY,
semanticCacheStore: semanticCache,
// Passing the cache does not switch it on; this asks for it on every request.
semanticCacheDefault: true,
cacheBreakpoints: "auto",
cacheBreakpointStore: {
get: (key) => kv.get(key),
set: (key, value, ttlMs) => kv.set(key, value, { px: ttlMs }),
},
});
// After handling requests with `layer`, save the calibration:
const save = async () => {
await semanticCache.settled();
await kv.set("bvsala:semantic-calibration", JSON.stringify(semanticCache.exportCalibration()));
};
return { layer, save };
}- Breakpoint placement. The store holds prompt-prefix fingerprints, timestamps and counts, one key per traffic class, never prompt text. The layer reads it before placing breakpoints and reads and rewrites it after the provider answers, so every instance extends the same measurement. If the store fails, that request is placed as a fresh instance would place it and nothing is written over what others learned. Without a store nothing changes.
- Calibration. The saved state is the whole experiment: both halves of the labels,
how many tests have already drawn on the confidence budget, and the floor in force.
Restoring it continues the same calibration, so the bound keeps holding across
restarts; it never starts the budget over. A state recorded under another embedder,
verifier, judge or configuration is refused with
CalibrationStateError. If you pass your ownembed,verifyorjudge, name them (embedderId,verifierId,calibration.judgeId) so a state can say what it was recorded with. - One line of history. The bound holds along one line: restore the latest state and
save from where it went. Instances that continue the same state at the same time
each spend its remaining confidence again, and the last one to save overwrites the
others' labels. For one bound across many concurrent instances, calibrate in a single
long-lived process, or offline with
calibrateSemanticFloor(), and give every instance the certified floor asacceptThreshold. The cached answers themselves still live in each instance's memory. - Why a predicted cache read did not happen (Anthropic's own API). With
cacheMissDiagnostics: true, the layer asks Anthropic to compare each marked request with the one whose cache entry it expected to read, and puts the provider'scache_miss_reasoninreport.traceswhen that read does not arrive. Anthropic keeps a short-lived fingerprint (hashes and token-count estimates) of each request that opts in; check that against your data-retention terms first.
Ask for a response contract and you also get data, validated against your schema:
const { data, text } = await layer.generate<{ answer: number }>({
model: "openai/gpt-4o-mini",
messages: [{ role: "user", content: "How many legs does a spider have?" }],
response: {
contractId: "answer:v1",
schema: {
type: "object",
properties: { answer: { type: "number" } },
required: ["answer"],
additionalProperties: false,
},
maxOutputTokens: 220,
},
});
console.log(data.answer); // 8A question the layer can answer on its own (the arithmetic of the first
example) is only served without the model when that answer meets your schema.
The arithmetic resolver returns { value: 8 }: with a { answer: number }
contract it steps aside and the request goes to the model, which answers in
your contract's shape.
For Gemini, use GoogleProvider: it speaks Gemini's own API (generateContent and
streamGenerateContent), where what the layer measures is documented. Keep
OpenAICompatibleProvider for every other provider.
import { GoogleProvider } from "@bivelio/savings-layer";
const gemini = createSavingsLayer({
provider: new GoogleProvider({ apiKey: process.env.GEMINI_API_KEY, requireUsage: true }),
licenseKey: process.env.BIVELIO_LICENSE_KEY,
});
const { text: answer, report: geminiReport } = await gemini.generate({
model: "google/gemini-3.8-flash", // the `google/` prefix never travels
messages: [
{ role: "system", content: "Answer in one sentence." },
{ role: "user", content: "Why is the sky blue?" },
],
maxOutputTokens: 400,
reasoningEffort: "low", // → thinkingLevel "LOW" (Gemini 3) or a thinking budget (2.5)
});
console.log(answer, geminiReport.cost.actual);What the native API gives the layer that the OpenAI-compatible endpoint hides:
- Thinking is billed output. Output is
candidatesTokenCount + thoughtsTokenCount(Google bills thinking as output) and the thinking is reported as reasoning. Thought summaries (thinkingConfig: { includeThoughts: true }on the adapter) are never part oftext. - Refusals with their reason. A blocked prompt (
promptFeedback.blockReason) or a filtered answer (finishReasonSAFETY,PROHIBITED_CONTENT,BLOCKLIST, …) isfinishReason: "content_filter"with the native reason infinishReasonDetail: measured from its usage, never cached, never credited. - Tools. Calls come back with their
idandthoughtSignature; send the assistant turn back with itstoolCallsas they came (as in the tool loop above) and the signature travels back untouched. The results of one turn go back together. - Implicit caching. Gemini 2.5 and later cache repeated prefixes on their own; the
cached part of the prompt (
cachedContentTokenCount) is reported as provider-cached tokens. Explicit context caching (cachedContents) is not used: it bills storage by the hour, which a per-request report cannot price. - Output contract. Your schema travels as written in
responseJsonSchema. Google does not document that every JSON Schema feature is enforced, so a repeat of a request with a schema is not valued by reference. - Thinking level.
reasoningEffortfollows Google's published mapping. A model that rejects a level (gemini-3.8-flashanswers 400 toMINIMAL) is asked again once atLOW, and remembered; the rejected request was not billed. Thinking cannot be turned off on Gemini 3 or 2.5 Pro, so"none"there is the lowest level, with a warning.
maxOutputTokens on the request caps the answer of any request, free text
included. With response.maxOutputTokens as well, the smaller of the two is sent.
const { text, finishReason } = await layer.generate({
model: "openai/gpt-5-mini",
messages: [{ role: "user", content: "Summarize this in two sentences: …" }],
maxOutputTokens: 300,
});On the wire it becomes Anthropic's max_tokens; on Chat Completions,
max_completion_tokens when baseUrl is OpenAI's own API (max_tokens is
deprecated there and its reasoning models, o-series and gpt-5.x, reject it) and
max_tokens on any other OpenAI-compatible endpoint, which is the field they all
understand. If an endpoint answers 400 naming the field it got as unsupported, the
request is repeated once with the other one (nothing was generated, so nothing is
billed twice) and the model is remembered. To choose the field yourself:
new OpenAICompatibleProvider({ baseUrl, maxTokensParam: "max_completion_tokens" }).
OpenAI-compatible endpoints that are not OpenAI (Gemini's
generativelanguage.googleapis.com/v1beta/openai, LiteLLM, vLLM, Ollama, a router).
prompt_cache_key, OpenAI's prefix-cache routing hint, goes only to OpenAI's own API:
Gemini answers 400 to a field it does not know, and up to 0.4.31 every request with a
cached prefix failed there. A proxy in front of OpenAI that forwards it can opt in with
sendPromptCacheKey: true. Any optional field an endpoint still rejects by name
(prompt_cache_key, service_tier, seed, reasoning_effort) is dropped, the request
repeated once and the model remembered; reasoning_effort is never dropped on OpenAI's
own API, where its 400 is the answer. Gemini leaves its thinking tokens out of
completion_tokens and bills them as output, so the layer counts output as
total_tokens − prompt_tokens whenever the total is larger (and reports the difference
as reasoning); on OpenAI's API completion_tokens already include reasoning and nothing
is added. Gemini stops a filtered answer with finish_reason: "content_filter: OTHER"
(or another reason): it is finishReason: "content_filter" (the adapter keeps the
reason, which reaches you as result.finishReasonDetail and is named in the report's
cache trace; an Anthropic refusal brings its stop_details.category there too), and like every
provider refusal it is measured, never
cached and never credited. Google does not document that such refusals are free, so
they are costed from their usage like any call.
The cap is yours, not a saving: it is never credited as output avoided. An answer
it cut (finishReason: "length") is not cached, and
report.qualityLedger.truncation says the limit was yours.
When you pass tools and the model asks to invoke one, the reply is not the final
answer: text is empty by design and the request is in toolCalls. Execute it,
append the assistant turn and one role: "tool" message per call to messages, and
call generate() again.
import type { ChatMessage, ToolDescriptor } from "@bivelio/savings-layer";
const tools: ToolDescriptor[] = [
{
id: "get_weather",
intents: ["weather", "forecast"],
risk: "read",
description: "Current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
additionalProperties: false,
},
},
];
// ← your tool
const getWeather = async (city: string): Promise<string> => `${city}: 21 °C, clear`;
const messages: ChatMessage[] = [{ role: "user", content: "What's the weather in Madrid?" }];
let res = await layer.generate({ model: "openai/gpt-4o-mini", messages, tools });
while (res.toolCalls?.length) {
// res.finishReason === "tool_calls". Send the calls back as they came, ids included.
messages.push({ role: "assistant", content: res.text, toolCalls: res.toolCalls });
for (const call of res.toolCalls) {
const args = JSON.parse(call.arguments) as { city: string };
const output = await getWeather(args.city);
messages.push({ role: "tool", toolCallId: call.id, name: call.name, content: output });
}
res = await layer.generate({ model: "openai/gpt-4o-mini", messages, tools });
}
console.log(res.text);- Ids. Each result carries the
toolCallIdof the call it answers, and every call of a turn gets its result before the next user or assistant message. On Anthropic that is what lets the turn travel astool_use/tool_resultblocks (with the tool still in that request'stools); otherwise it is sent as plain text, as before, and the SDK warns once. - Errors. When a tool fails, put the error in
contentand setisError: true. Anthropic receives it asis_error; providers without such a flag get the text. - Gemini 3's thought signature. With
GoogleProviderit travels in each call'sthoughtSignatureand comes back incall.extraContentin the same shape as below, so a conversation can move between the two Gemini adapters. Through its OpenAI-compatible endpoint Gemini returns each call withextra_content.google.thought_signatureand rejects the next turn with a 400 if the call comes back without it. It arrives incall.extraContent; pushingres.toolCallsback as they came (as above) sends it back untouched. It is opaque and not part of any cache key. - gpt-5.6-sol, gpt-5.6-terra and gpt-5.6-luna with tools need
reasoningEffort: "none". On Chat Completions OpenAI rejects function tools for these models unlessreasoning_effortis"none"(«Function tools with reasoning_effort are not supported for gpt-5.6-sol in /v1/chat/completions»). Declare it on the request:layer.generate({ model: "openai/gpt-5.6-sol", messages, tools, reasoningEffort: "none" }). The layer never picks an effort itself: without the field nothing is sent. It is part of the cache key, liketemperatureandseed.
A tool‑call turn is never served from cache and never counted as an avoided call:
an action that has not run yet cannot be a saving. The answer that follows a tool loop
can be, and its cache key covers the loop: each call's name and arguments, which call
each result answers and its isError. Two conversations that differ only there never
share a cached answer; the fresh call ids of a new run of the same loop are not a
difference.
Which tools reach the model. The layer never cuts a catalogue it cannot rank. If
your tools declare no intents and no risk (the usual shape of MCP and OpenAI tool
lists), or you pass savings: { router: false }, every tool goes to the provider
exactly as you sent it, and the layer credits nothing for the tool block. When your
tools declare intents/risk, the selector holds back write tools from a request
that is not an action and, since 0.4.33, sends every other tool: it prunes by
relevance only if you ask for it with new ToolSelector({ maxTools }) (up to 0.4.32
that was the default, with 8). Whatever it leaves out, the model cannot call in that
request: the report's tool-selector trace lists it, the tokens count as eliminated
input, and the SDK warns once per process (silenceWarnings turns the warning off).
Large catalogues: let the provider search them (opt‑in). With hundreds of tools
(several MCP servers, say), savings: { toolSearch: "auto" } sends every tool deferred
behind the provider's own tool search instead of putting the whole catalogue in every
request's prompt. The model looks up what it needs, no tool you declared is out of
reach, and the cached prompt prefix stays the same from one question to the next.
- Where. The Anthropic adapter, on models in Anthropic's tool search compatibility table. OpenAI documents its tool search for the Responses and Agents APIs, not for Chat Completions, which is what this SDK's OpenAI adapter speaks, so OpenAI requests keep the selector.
- When.
"auto"applies it when the adapter talks to Anthropic's own API, the catalogue is over the provider's documented size, and deferring it pays for one search on every request even with a warm prompt cache."always"skips those checks (for example behind a proxy that forwards the search blocks). Requests with an output contract, or with the model cascade on, keep the selector. - Across turns. Append the assistant turn with its
toolCalls(ids as returned) before the tool results, as in any tool conversation. The Anthropic adapter then sends the earlier search results back in that turn, so the model calls the tools it already found without searching again. It keeps them in memory, per adapter instance and bounded; a turn it no longer remembers (another process, say) goes out as before and the model searches again. - Whose saving. The provider's. Your usage counter shows it; the layer credits nothing for it and charges no commission on it.
import { AnthropicProvider } from "@bivelio/savings-layer";
const claude = createSavingsLayer({
provider: new AnthropicProvider({ apiKey: process.env.ANTHROPIC_API_KEY }),
licenseKey: process.env.BIVELIO_LICENSE_KEY,
});
const mcpTools: ToolDescriptor[] = []; // every tool your MCP servers list
const { report: searched } = await claude.generate({
model: "anthropic/claude-opus-5-5",
messages: [{ role: "user", content: "Which orders from service 17 are still open?" }],
tools: mcpTools,
savings: { toolSearch: { mode: "auto", alwaysLoaded: ["read_file"] } },
});
// Which path ran and why, and how many searches the provider ran.
console.log(searched.traces.filter((t) => t.stage === "tool-search"));Wrap the layer once and every request becomes a span on the OpenTelemetry pipeline
you already run (@opentelemetry/api is an optional peer; the SDK never imports it):
import { trace } from "@opentelemetry/api";
import { instrumentSavingsLayer } from "@bivelio/savings-layer";
// One span per request, on the OpenTelemetry pipeline you already run.
const traced = instrumentSavingsLayer(layer, { tracer: trace });
const { report: tracedReport } = await traced.generate({
model: "openai/gpt-4o-mini",
messages: [{ role: "user", content: "What is 2 + 2 * 3?" }],
});
console.log(tracedReport.cost.basis); // the span carries the same figures- Standard spans, plus your ledgers. Spans follow the
GenAI semantic conventions
as last published (OpenTelemetry semantic conventions v1.41.1):
chat {model}, kind CLIENT, provider, model and usage from your provider's counter (0 on an avoided call: no token was billed). On top,com.bivelio.*carries the four ledgers, thebasis, the cost and the ids;com.bivelio.request_idis the same id as your usage record. Same results,
