@metrxbot/sdk
v0.5.0
Published
TypeScript SDK for integrating with Metrx gateway
Readme
Metrx SDK
TypeScript SDK for integrating with the Metrx gateway. The gateway is a Cloudflare Worker that proxies LLM API calls with cost tracking, rate limiting, and request metadata.
Installation
npm install @metrxbot/sdkQuick Start
Initialize the Client
import { MetrxbotClient } from '@metrxbot/sdk';
const client = new MetrxbotClient({
apiKey: process.env.METRX_API_KEY,
gatewayUrl: 'https://gateway.metrxbot.com', // optional, defaults to this
defaultAgentId: 'my-agent', // optional
timeout: 30000, // optional, in milliseconds
});Chat Completions (OpenAI-compatible)
const response = await client.chat({
model: 'gpt-4',
messages: [{ role: 'user', content: 'What is 2+2?' }],
providerKey: process.env.OPENAI_API_KEY, // REQUIRED: your provider key — the
// gateway forwards it to the provider (your Metrx key only authenticates
// the gateway; without providerKey the call fails "Missing X-Provider-Key")
customerId: 'user-123', // optional, for tracking
sessionId: 'session-456', // optional, for tracking
});
console.log(response.choices[0].message.content);
console.log(`Cost: ${response._meta.costMicrocents} microcents`);
console.log(`Latency: ${response._meta.latencyMs}ms`);Streaming Chat Completions
const stream = client.chatStream({
model: 'gpt-4',
messages: [{ role: 'user', content: 'Write a haiku' }],
});
for await (const chunk of stream) {
if (chunk.choices?.[0]?.delta?.content) {
process.stdout.write(chunk.choices[0].delta.content);
}
}Anthropic Messages API
const response = await client.messages({
model: 'claude-3-sonnet',
max_tokens: 1024,
messages: [{ role: 'user', content: 'Hello, Claude!' }],
});
console.log(response.content[0].text);Embeddings
const response = await client.embeddings({
model: 'text-embedding-3-small',
input: 'Hello world',
});
console.log(response.data[0].embedding);Health Check
const health = await client.health();
console.log(health.status); // 'ok', 'degraded', or 'error'
console.log(health.version);List Available Models
const models = await client.models();
console.log(models.data.map((m) => m.id));Configuration
Client Options
interface MetrxbotConfig {
// Required: Your Metrx API key
apiKey: string;
// Optional: Gateway URL (defaults to https://gateway.metrxbot.com)
gatewayUrl?: string;
// Optional: Default agent ID for all requests
defaultAgentId?: string;
// Optional: Default LLM provider (openai, anthropic, etc.)
defaultProvider?: string;
// Optional: Request timeout in milliseconds (defaults to 30000)
timeout?: number;
}Request Options
All API methods accept these additional optional parameters:
interface RequestOptions {
// Override the default agent ID for this request
agentId?: string;
// Customer/end-user ID for tracking
customerId?: string;
// Session ID for tracking
sessionId?: string;
// Provider API key (if required by provider)
providerKey?: string;
// Force a specific provider for this request
provider?: string;
}Error Handling
The SDK provides specific error classes for different failure scenarios:
import {
MetrxbotError,
AuthenticationError,
RateLimitError,
GatewayError,
ValidationError,
TimeoutError,
} from '@metrxbot/sdk';
try {
const response = await client.chat({
model: 'gpt-4',
messages: [{ role: 'user', content: 'Hello' }],
});
} catch (error) {
if (error instanceof AuthenticationError) {
console.error('API key is invalid');
} else if (error instanceof RateLimitError) {
console.error(`Rate limited. Retry after ${error.retryAfter}s`);
} else if (error instanceof GatewayError) {
console.error('Gateway is experiencing issues');
} else if (error instanceof ValidationError) {
console.error('Invalid request parameters');
} else if (error instanceof TimeoutError) {
console.error('Request timed out');
} else {
console.error('Unknown error:', error);
}
}Response Metadata
All responses include a _meta field with gateway metadata:
interface MetrxbotMeta {
// Gateway response latency in milliseconds
latencyMs: number;
// Cost of the request in microcents
costMicrocents: number;
// X-Request-ID header from gateway
requestId?: string;
}Supported Environments
- Node.js 18+
- Deno
- Modern browsers (with native fetch support)
Cascade routing (v0.2.1, opt-in)
Delegate per-request model choice to Metrx inside your own process — your provider keys never leave your environment, and prompt/output text is never sent to Metrx (scores only).
import { MetrxbotClient, runWithPolicy } from '@metrxbot/sdk';
const metrx = new MetrxbotClient({ apiKey: process.env.METRX_API_KEY! });
const r = await runWithPolicy(metrx, {
agentKey: 'ask_fundry',
input: prompt,
// Your own provider call — your client, your keys.
// model === undefined ⇒ call your own default (day-0 pass-through).
callModel: (model, input) => callAnthropic(model ?? DEFAULT_MODEL, input),
// Cross-family checker (required before shadow/cascade modes are enabled):
checkerCaller: (model, task) => callOpenAI(model!, task),
});
use(r.output);
// Attach r.requestId to your telemetry events (request_id field) so cost
// joins work: it links llm_events ↔ optimization_decisions ↔ quality.Behavior is controlled server-side per agent (routing_policies): with no
policy row every call is an exact pass-through. Modes advance
pass_through → shadow_sample → cascade, each gated by evidence. Every
seam fails open — no code path throws an error your own
callModel(prodModel) would not have thrown. Kill switches: server policy
enabled=false, env METRX_ROUTING_DISABLED=true, min_sdk_version.
Not supported in v1: streaming call sites (keep them on their original path).
Tagging telemetry via ctx (v0.2.1)
Both callModel and checkerCaller receive an optional third argument, a
ModelCallContext, on every provider call the router makes. It's purely
informational — the router serves exactly the same thing whether or not you
read it — but it lets your own instrumentation stamp routing tags on the
provider call so cascade cost SUMs join across
llm_events ↔ optimization_decisions ↔ quality, and skip the background
shadow candidate so it never lands on your cost dashboard.
import { runWithPolicy, type ModelCallContext } from '@metrxbot/sdk';
const r = await runWithPolicy(metrx, {
agentKey: 'ask_fundry',
input: prompt,
callModel: (model, input, ctx?: ModelCallContext) => {
// Don't record the background shadow candidate — it's never served and
// would double-count on your cost dashboard.
const record = ctx?.phase !== 'shadow_candidate';
return callAnthropic(model ?? DEFAULT_MODEL, input, {
// Stamp the same request_id the decision/quality rows carry, plus the
// experiment id + arm, so your cost query joins on all three.
tags: record
? {
request_id: ctx?.requestId,
experiment_id: ctx?.experimentId,
arm: ctx?.arm, // 'control' | 'treatment'
phase: ctx?.phase, // 'primary' | 'candidate' | 'escalation' | 'checker' | 'shadow_candidate'
}
: undefined,
});
},
checkerCaller: (model, task, ctx) => callOpenAI(model!, task), // ctx.phase === 'checker'
});ctx is optional and additive: existing two-arg callModel/checkerCaller
closures keep working unchanged (the extra argument is simply ignored).
ctx.requestId always equals RunWithPolicyResult.requestId.
Serverless cold start (v0.2.3)
The policy cache is lazy: the first run() on a fresh instance returns
pass-through while the policy set loads in the background. On long-lived servers
this is invisible. On serverless (Vercel/Lambda), where most invocations are
cold at low volume, it means the first call per instance per agent never enters
shadow/cascade — so a low-traffic experiment can be starved of decisions. Two
opt-in fixes (both fail-open, both no-ops on a warm cache):
import { MetrxbotClient, routerFor } from '@metrxbot/sdk';
const metrx = new MetrxbotClient({ apiKey: process.env.METRX_API_KEY! });
// Preferred: warm once at module init — zero request-path latency.
const router = routerFor(metrx);
await router.warm(); // bounded (default 5s), never throws
// OR: let the first call await the load itself (adds a one-time ~policy-fetch
// latency to the first call per instance; bounded by coldCacheAwaitTimeoutMs).
const router2 = routerFor(metrx, { awaitPolicyOnColdCache: true });Both default OFF — omit them and behavior is byte-identical to prior versions.
Use warm() if you have a natural init point; use awaitPolicyOnColdCache if
you don't. (Emit POSTs also set keepalive as of 0.2.3, so decision/pair
telemetry is more likely to survive a serverless freeze-on-response.)
router.flush() is REQUIRED on serverless
The shadow leg (candidate call + judging + the pair POST) runs in the
background after your response is served. On serverless runtimes
(Vercel/Lambda), the instance is frozen the moment your handler returns —
which kills that background work mid-flight. The observable symptom: decisions
land (they fire on the serving path with keepalive), but judged pairs never
arrive. This exact omission cost one integration five weeks of debugging
(ISSUE-552).
const result = await runWithPolicy(metrx, { agentKey, input, callModel, checkerCaller });
// ... use result.output ...
await router.flush(); // REQUIRED before returning on serverlessOn long-lived servers flush() is optional (call it at graceful shutdown).
Run npx @metrxbot/sdk doctor --no-flush to see the failure mode reproduced.
Automatic event↔decision join (v0.5.0)
runWithPolicy() now stamps its requestId onto OpenTelemetry spans created
inside serving-path model calls (primary / served candidate / escalation)
via MetrxbotSpanProcessor, and MetrxbotSpanExporter sends it as the event's
request_id — which Metrx stores as llm_events.gateway_request_id, the join
key for optimization_decisions.request_id. The manual ctx.requestId
tagging above still works and always wins when you set it explicitly.
Scope caveats: shadow_candidate and checker legs are deliberately NOT
stamped (diagnostic spend must not pollute per-request cost joins), and
telemetry-only call sites that never go through runWithPolicy() have no
request to join — they stay correlated by task_id only.
Diagnostics: metrx doctor (v0.5.0)
METRXBOT_API_KEY=sk-... npx @metrxbot/sdk doctor --agent-key my_agentTraces one shadow leg hop-by-hop — policy fetch → decision emit → candidate → judge → pair POST → event export — and prints PASS/FAIL per hop with the exact failing hop and its remediation. Zero-write: the leg runs on an in-process capture transport; real network is limited to the read-only policy GET and empty-batch reachability probes that the API rejects before any insert.
Flags: --agent-key (diagnose your real policy), --no-flush (reproduce the
serverless no-flush failure), --no-network (offline), --json.
License
MIT
