@agentoria/ratecard
v0.2.1
Published
LLM pricing data and a tier-aware cost calculator. Bundled offline snapshot, optional live refresh.
Maintainers
Readme
@agentoria/ratecard
LLM pricing data with a tier-aware cost calculator. Zero dependencies, works offline, optionally refreshes from the published API.
npm install @agentoria/ratecardWhy
Most pricing tables are a flat $/1M number per model. Real bills are not
flat: context ladders change the rate mid-call, batch APIs halve it, prompt
caching discounts part of the input, and half the vendors publish in CNY. This
package encodes all of that and tells you which rule it applied.
Quick start
import { price } from '@agentoria/ratecard';
const r = price('gpt-5.5', { inputTokens: 300_000, outputTokens: 2_000 });
r.total // 3.09
r.currency // 'USD'
r.tier.index // 1 — the call crossed the 272K boundary
r.tier.minTokens // 272001
r.breakdown // { input: 3, cachedInput: 0, output: 0.09 }
r.unitPrices // { input: 10, cachedInput: 1, output: 45 }
r.warnings // []Model ids accept vendor aliases, so you can pass the same string you passed to the vendor's SDK:
price('deepseek-chat', { inputTokens: 1_000, outputTokens: 500 });Prompt caching
Pass the absolute cached-token count — the number vendors return in their usage payload — not a ratio:
price('gpt-5.5', {
inputTokens: 100_000,
outputTokens: 2_000,
cachedInputTokens: 80_000, // billed at the cache-hit rate
});If the model publishes no cached-input price, the tokens are billed in full and you get a warning saying so. Nothing is discounted silently.
Batch API
price('qwen3-max', usage, { mode: 'batch' });Models without published batch rates fall back to standard pricing, with a
warning. Check result.warnings — it is the difference between "this is the
batch price" and "this is what you'd pay if batch existed".
Currencies
Prices are stored in each vendor's own pricing-document currency. Ask for whichever you want:
price('deepseek-chat', usage); // CNY — DeepSeek's native
price('deepseek-chat', usage, { currency: 'USD' }); // convertedThe bundled FX table travels with the data. Override it when you need a specific rate:
price('gpt-5.5', usage, {
currency: 'CNY',
fxRates: { base: 'USD', date: '2026-07-01', rates: { USD: 1, CNY: 7.1 } },
});Comparing models
import { compare } from '@agentoria/ratecard';
const { results } = compare(
{ inputTokens: 10_000, outputTokens: 1_000 },
{ filter: { country: 'CN', capabilities: ['tool-use'] } },
);
results[0]; // cheapest, priced in USD so the ranking is meaningfulcompare normalises to USD by default — ranking a CNY total against a USD
total would silently compare different units.
Browsing the catalogue
import { listModels, getModel, models, providers } from '@agentoria/ratecard';
listModels({ providerId: 'anthropic' });
listModels({ minContextLength: 200_000, status: 'ga' });
listModels({ search: 'flash' });
getModel('claude-opus-4-7')?.source.url; // the vendor page it came fromEvery record carries source.url and source.fetchedAt — the official page it
was transcribed from, and when. Prices change without notice; check the source
before making a procurement decision on a number.
Live data
The package ships a snapshot, so everything above works with no network. To pick up newer prices without upgrading the package:
import { Client } from '@agentoria/ratecard';
const client = new Client({ ttlSeconds: 3600 });
await client.refresh(); // never throws
client.dataset.price('gpt-5.5', usage);refresh() reports failures in its return value instead of throwing, and keeps
serving the data it already has. A CDN blip degrades you to slightly stale
prices; it does not take down your request path.
const r = await client.refresh();
if (r.error) console.warn('using cached prices:', r.error.message);It fetches a ~1 KB manifest first and only downloads the catalogue when
dataVersion actually changed.
Estimating tokens
import { estimateTokens } from '@agentoria/ratecard';
estimateTokens('你好,世界'); // heuristic, ±10–20% of a real tokeniserDeliberately not tiktoken: that is megabytes of WASM, matches one vendor, and
this package has to stay dependency-free. Use it for budgeting; use the
vendor's reported token counts for anything that has to reconcile.
When the estimate is not good enough, supply a real tokeniser instead of replacing the module — a Node-side adapter is one line and stays out of the browser bundle:
import { encode } from 'gpt-tokenizer';
import { Dataset, bundledSnapshot } from '@agentoria/ratecard';
const dataset = new Dataset(bundledSnapshot, { tokenizer: (t) => encode(t).length });
const r = dataset.priceText('gpt-5.5', prompt, 500);
r.estimatedInputTokens; // the count actually used
r.tokenizerUsed; // 'custom' — so you can tell an estimate from a measurementBillable amounts
total is a float, and that is fine: measured across all 3,204 conformance
cases the worst relative error against exact decimal is 2.8e-16 — one unit in
the last place. What is not fine is scaling it to cents yourself:
Math.round(0.025 * 100); // 3
toMinorUnits(0.025, 'USD'); // 2n — exact, rounds half to evenimport { toMinorUnits, toBillableAmount } from '@agentoria/ratecard';
const r = price('gpt-5.5', usage);
toMinorUnits(r.total, r.currency); // 309n — the integer you invoice
toBillableAmount(r.total, r.currency); // 3.09The conversion goes through the value's decimal representation with integer arithmetic, never a float multiply, and the Python and Go SDKs quantize identically.
Errors
Invalid usage throws rather than returning a plausible-looking number:
price('gpt-5.5', { inputTokens: -1, outputTokens: 0 }); // RangeError
price('gpt-5.5', { inputTokens: 10, outputTokens: 0,
cachedInputTokens: 11 }); // RangeError
price('nope', usage); // ErrorTier semantics
Each tier states what its token range measures, because vendors differ:
| basis | range is compared against |
|---|---|
| input_tokens | prompt tokens (default, most common) |
| total_tokens | prompt + completion |
| context_window | the model's declared window, fixed per deployment |
result.tier tells you which tier was selected and on what basis, so a
surprising number is always traceable.
Data source and license
Data is transcribed from official vendor pricing pages by
LLM Ratecard and published at
/api/v1/. Code is MIT; the dataset is CC BY 4.0 and asks for attribution:
Pricing data from LLM Ratecard (https://github.com/WangYihang/llm-ratecard), CC BY 4.0.
It is a community transcription, provided without warranty of accuracy. Verify against the linked official page before relying on a number for billing.
Cross-language parity
The pricing algorithm is pinned by conformance/cases.json in the repo — 3,200
cases covering every model. The Python and Go SDKs are held to the same
fixture, so all implementations produce identical numbers.
