@usageflow/vibe
v0.2.1
Published
(Beta) Unified Anthropic + OpenAI client wrapper with inline UsageFlow sync usage reporting, plus manual withdraw()/credit() ledger adjustments
Maintainers
Readme
@usageflow/vibe (beta)
Beta:
@usageflow/vibeis under active development. APIs may change between minor versions — pin a version if that matters to you, and please report issues on GitHub.
@usageflow/vibe is a single, framework-agnostic TypeScript library that wraps the Anthropic and OpenAI SDKs behind one interface, gating every call on an inline UsageFlow sync usage report.
Usage is recorded right after each call, not necessarily before it returns — so credits() can lag a moment behind a call that just resolved. withdraw(), credit(), and close() resolve once the request has been sent, not once it's been applied.
Install
npm install @usageflow/vibe
export USAGEFLOW_API_KEY='your-api-key'Usage
import { usage } from '@usageflow/vibe';
usage.init({ apiKey: process.env.USAGEFLOW_API_KEY! });
const result = await usage.chat({
identity: 'cust_acme',
workflow: 'support-agent', // optional — matched against a Vibe policy's slug, not the metering key
provider: 'anthropic',
model: 'claude-sonnet-5',
messages: [{ role: 'user', content: 'Summarize this ticket.' }],
});
console.log(result.content);identity is the ledger key everything is metered/attributed against — it's sent as alias on the wire. workflow is optional and sent as allocationMetadata.workflowId; the server matches it against a Vibe policy's slug (a policy named "Downgrade OpenAI Models" gets slug downgrade-openai-models) to apply that policy's workflow-scoped rules. model is always sent too (allocationMetadata.model) — the server reads it back as requestedModel for the policy's audit log and lets a tier's condition branch on it. A workflow with no matching policy is simply ignored server-side — it does not fail the call.
Vibe policy tiers can redirect the call
If workflow resolves to an active policy and a tier's condition matches, the server can hand back a decision instead of a plain approval — and chat() follows it automatically:
ROUTE_MODEL/DEGRADEwith amodelparam —chat()dispatches to that model instead of the one you requested, inferring the provider from the model id (claude-*→ Anthropic, everything else → OpenAI), so a policy can drop an Anthropic call down to an OpenAI model or vice versa.result.providerandresult.modelreflect what actually ran, andresult.vibePolicycarries the full decision ({ policyId, tierIndex, action }) for logging.BLOCK— never reaches you as a decision object; it's the same denial path as running out of quota, surfaced asVibeRejectionError.- A tier action with no
modelparam —result.vibePolicyis still populated (so you can see a tier fired, e.g. for a Slack/webhook-only effect), but the original model/provider is used untouched.paramsis free-form exactly as configured in the console's tier editor — treat unrecognized keys as forward-compatible extras.
const result = await usage.chat({
identity: 'cust_acme',
workflow: 'downgrade-openai-models',
provider: 'anthropic',
model: 'claude-sonnet-5',
messages: [{ role: 'user', content: 'Summarize this ticket.' }],
});
if (result.vibePolicy) {
console.log(`Vibe policy ${result.vibePolicy.policyId} tier ${result.vibePolicy.tierIndex} fired`);
// if it redirected the model, result.provider / result.model already reflect that —
// no extra handling needed to actually use the decision.
}Provider keys
Vibe uses your own Anthropic and OpenAI keys. If ANTHROPIC_API_KEY / OPENAI_API_KEY are set in the environment — the same variables the official provider SDKs read — they are picked up automatically; you only need the one for the provider you call. You can always pass them in init() instead, which takes precedence:
usage.init({
apiKey: process.env.USAGEFLOW_API_KEY!,
anthropicApiKey: '...',
openaiApiKey: '...',
});OpenAI reasoning models (o1/o3/o4/gpt-5)
Handled automatically: maxTokens is sent the way these models expect, and temperature is dropped because they only accept the default.
Attach customer metadata
Pass business context (plan, tier, region, department, etc.) that UsageFlow does not and cannot infer on its own:
const result = await usage.chat({
identity: 'cust_acme',
provider: 'anthropic',
model: 'claude-sonnet-5',
messages: [{ role: 'user', content: 'Summarize this ticket.' }],
customerMetadata: { plan: 'pro', region: 'eu' },
});customerMetadatais durable context about whoidentityis — the backend persists it against that identity, so it applies to that identity's future calls too, not just this one. There is no separate per-call-only metadata concept in Vibe — everything you pass here is durable.- Only string, number, and boolean values are supported; other value types are dropped.
- Sent as its own field on the wire, kept separate from
workflow'sallocationMetadata.workflowId(workflow-slug policy routing) and from the free-formmetadatafield (trace annotations) — none of these are merged together.
withdraw()/credit() accept the same customerMetadata field.
Manual ledger adjustments: withdraw() / credit()
For usage that happens outside a metered chat() call — background jobs, batch corrections — deduct or reverse against an identity directly:
const result = await usage.withdraw({
identity: 'cust_acme',
amount: 500,
unit: 'credits', // default, descriptive only
idempotencyKey: 'job-42-attempt-1',
reason: 'nightly batch reconciliation',
});
console.log(result.eventId); // ledger event id, for audit correlation
// Correcting an over-count:
await usage.credit({
identity: 'cust_acme',
amount: 500,
idempotencyKey: 'job-42-reversal-1',
reason: 'correct over-count from job-42',
});There's no separate "withdraw" concept on the server — only request_for_allocation (policy/quota check) and use_allocation (settle), the same pair chat() uses. withdraw()/credit() compose both client-side, back to back: the sync request_for_allocation call is what actually decides why a withdrawal is allowed or denied (server-side policy, same mechanism as chat()), and use_allocation immediately settles the same amount — since a manual withdrawal's amount is already final, unlike chat()'s estimate-then-real-usage split. credit() sends the same request as a negative amount through that same pair.
Both the chat() settlement and the withdraw()/credit() settlement are sent with waitForConfirmation: true and awaited, so the server settles in Postgres before ACKing rather than taking the default Kafka-only fast path — that fast path depends on a separate consumer to materialize the settlement, and a failure on it is otherwise invisible to the caller. If settlement fails after an approved withdraw()/credit() reservation, the call throws VibeRejectionError — the deduction never happened. chat() settlement failures are logged, not thrown, since by the time settlement runs the provider call has already completed and its result is real.
If the policy check denies the reservation, no settlement is attempted and the call throws VibeRejectionError (same error chat() throws) — no partial deduction. There is currently no balance in the result: the server only returns an allocationId (exposed as eventId), not a running balance. idempotencyKey is required on the request and carried in the ledger event's metadata for audit/correlation, but the server does not yet deduplicate on it — a literal retry can still double-deduct until server-side dedup ships. credit()'s negative-amount behavior is unverified against server policy (built for consumption, not refunds) — confirm before relying on it in production.
Deduct now, settle later: withdrawAsync() / creditAsync()
For work where you don't know the final amount up front — start a job, find out what it cost once it finishes — reserve the deduction now and settle it later:
const capture = await usage.withdrawAsync({
identity: 'cust_acme',
amount: 500, // worst-case estimate to reserve
idempotencyKey: 'job-42-attempt-1',
// holdForMs: 60 * 60 * 1000, // optional — how long the hold stays open; default 24h
});
console.log(capture.captureId); // pass to close() once the job finishes
console.log(capture.expiresAt); // epoch ms — the hold must be closed before this
// ... later, once the real cost is known:
await usage.close({ captureId: capture.captureId, amount: 420 }); // settles the final amountcreditAsync() is the async counterpart of credit() — same shape, reserves a negative amount, settled the same way with close().
The hold defaults to 24 hours and can be overridden per call with holdForMs (milliseconds, must be greater than zero). If close() never runs before the hold expires, the hold is released automatically and nothing is charged — the reserved amount is never deducted on its own. close()'s amount defaults to the originally reserved amount if omitted, and should not exceed it. A capture can only be closed once; identity on close() is only needed when it runs in a different process than the one that created the capture.
Show users their credits: credits()
Read-only — nothing is reserved or charged. Call it from your backend to show a user what they have left:
const c = await usage.credits({
identity: 'cust_acme',
workflow: 'free-landing-page', // optional — omit for every active workflow
});
c.account; // plan request cap: { used, limit, remaining, periodEndAt, exceeded } (limit null = unlimited)
c.identity; // { identity, used, known } — known:false = never made a request
c.workflows[0].remaining; // credits left before this workflow blocks (null when it never blocks)
c.workflows[0].blocked; // true once the limit is reached
c.workflows[0].resetInterval; // e.g. "1d" — how often usage renewsUsage is counted per identity, not per workflow, so every workflow's limit is measured against the identity's one used number. customerMetadata (e.g. { plan: 'pro' }) can be passed to pick a policy branch for an identity with no stored metadata. Against a UsageFlow server that predates get_credits, credits() throws VibeCreditsUnsupportedError.
Behavior (v0)
usage.chat({ provider, ... })resolves and constructs the matching Anthropic/OpenAI SDK client internally, lazily, the first time that provider is used — no client instances to construct or pass in yourself.- Every
usage.chat()call hits UsageFlow's sync usage-report event first and awaits the response before dispatching to the provider SDK. - The pre-flight reservation is a real worst-case ceiling — an estimated input token count plus the enforced
maxTokensoutput cap (default 1024) — not a placeholder, so over-budget calls can actually be rejected up front. - A denial (over quota/policy) surfaces as
VibeRejectionError— the provider call never happens. - No remote config fetching on the request path; policy overrides (model downgrade, effort caps) are out of scope for this incarnation.
- After the call completes, UsageFlow is settled with real token usage (never more than the reservation), not the reserved ceiling.
Production notes
Depends on @usageflow/core for the underlying WebSocket/allocation transport (pooled connections, reused across calls), and on @anthropic-ai/sdk + openai directly — both ship as real dependencies since vibe owns their instantiation.
Node.js 18+ is recommended. MIT licensed. See the npm package or repository.
