@autter/runtime-node
v1.5.0
Published
Autter Runtime for Node.js: request wide events, coded errors, operation logging, same-origin browser relay and a curated OpenTelemetry server tracker
Maintainers
Readme
Operation logging
Use withRuntimeOperation, runtimeLogger, and createRuntimeLogger to attach
application context to logs and completed operations. See the full
operation logging guide for initialization,
steps, outcomes, redaction, export lifecycle, and release requirements.
Requests and coded errors (1.5.0)
Requires ingester 1.5.0. One summary per request (always kept), request ids echoed to clients, one issue per error code:
import {
autterRequests, autterErrorResponse, defineRuntimeErrors, runtimeContext,
} from "@autter/runtime-node";
export const billingErrors = defineRuntimeErrors("billing", {
declined: { status: 402, message: "Payment declined", expected: true,
why: "The card issuer rejected the charge", fix: "Ask for another card" },
limit: ({ plan }: { plan: string }) => ({ status: 429, message: `Plan ${plan} limit reached` }),
});
app.use(autterRequests({ ignore: ["/healthz", "/metrics"] }));
app.post("/checkout", async (req, res) => {
runtimeContext.set({ cart: { items: req.body.items.length } });
runtimeContext.info("Applied coupon", { coupon: "SPRING" }); // folded into the summary
if (overLimit) throw billingErrors.limit({ plan: "free" }); // Express 5 (next(err) on 4)
res.json({ ok: true, requestId: runtimeContext.requestId });
});
app.use(autterErrorResponse()); // { error: { message, code, why?, fix?, link?, requestId } }Also: autterFastify, withRuntimeRequest (fetch-style handlers),
runtimeContext.fork / runInBackground / carrier() for background work and
queues, the per-operation AI usage rollup, waitUntil, initAutterLogging
(no NodeSDK — for apps with their own OTel), enrichers (enrichUserAgent,
enrichRequestSize, enrichEdgeGeo, enrichDeployment), sinks (otlpSink,
consoleSink, fileSink → .autter/runtime/*.jsonl in development) and test
helpers:
import { captureRuntime, expectOperation } from "@autter/runtime-node/testing";
const runtime = captureRuntime();
// … exercise the app …
expectOperation(runtime, "POST /checkout").toHaveOutcome("degraded").toHaveErrorCode("billing.declined");Full guide: Requests, coded errors and background work.
@autter/runtime-node
The server tracker exports per-instance RSS, heap, memory limit, and GC
metrics by default. See memory pressure incidents
for ECS/Kubernetes OOM event forwarding and optional heap profiles. Set
memoryMetrics: false only if another exporter already emits equivalent
process metrics.
Autter Runtime for Node.js — two halves in one package:
- Same-origin browser relay for
@autter/runtime-browser - Curated OpenTelemetry server tracker exporting OTLP/HTTP JSON
Install
npm install @autter/runtime-node1. Browser relay
The browser tracker posts to your backend; this handler validates and whitelist-sanitises the payload, attaches your private ingest key server-side, forwards asynchronously, and returns 202 immediately.
Express / Node http:
import { createBrowserRelayHandler } from "@autter/runtime-node";
app.post(
"/api/autter-runtime",
createBrowserRelayHandler({ apiKey: process.env.AUTTER_RUNTIME_KEY! }),
);Next.js App Router / any fetch-style runtime:
import { createBrowserRelayFetchHandler } from "@autter/runtime-node";
export const POST = createBrowserRelayFetchHandler({
apiKey: process.env.AUTTER_RUNTIME_KEY!,
});Options: endpoint (default https://otlp.autter.dev), maxBodyBytes
(default 64 KB), onError.
2. Server tracker
// instrument.ts — must run before anything else creates connections
import { initAutterServer } from "@autter/runtime-node";
const autter = initAutterServer({
apiKey: process.env.AUTTER_RUNTIME_KEY!,
service: "payments-api",
environment: process.env.NODE_ENV,
release: process.env.GIT_SHA,
});
// handled errors — always recorded, never sampled out:
autter.captureException(err, { "order.id": "…" });
// warnings/info without an exception — same grouping+aggregation as
// errors, just a lower severity ("fatal" | "error" | "warning" | "info"):
autter.captureMessage("Legacy /orders lookup used", "warning");
// named process spans — background jobs, queue consumers, cron ticks,
// DB-heavy calls. Always recorded (never head-sampled), so Autter's
// slow-process monitor sees accurate run counts and durations, and can
// flag the process when it is slow and repeating a lot:
await autter.withProcessSpan("invoice.rebuild", async () => {
await rebuildInvoices();
});
// graceful shutdown flushes exporters:
await autter.shutdown();Run it first: node --require ./instrument.cjs server.js (CJS), or for
pure-ESM apps add OTel's loader hook
(node --import ./instrument.mjs --experimental-loader=@opentelemetry/instrumentation/hook.mjs server.js)
so http auto-instrumentation can patch ESM imports.
Defaults (cheap by construction):
| Signal | Default |
| --- | --- |
| Captured/unhandled exceptions | 100% (dedicated always-on tracer) |
| Traces containing an error | 100% (tail retention, retainTracesOnError) |
| withProcessSpan spans | 100% (same always-on tracer) |
| LLM/GenAI call spans | 100% (llmTracing, on by default) |
| Healthy traces | 1% head sampling (traceSampleRate) |
| Request metrics | exported every 60 s |
| Logs | not collected |
| PII in custom attributes | redacted before export (redactAttributes) |
| Forgotten shutdown() | exporters still flushed on exit (autoFlush) |
Error-linked trace retention. Errors export at 100% while traces are
head-sampled — on its own that strands a retained error without the trace
that explains it. So unsampled spans are kept briefly in an in-process
buffer, and the moment a trace shows an error — a 5xx response, a recorded
exception, captureException, or an error-severity captureMessage — the
whole trace is exported, sampling lottery notwithstanding. The buffer is
bounded (256 spans per trace, 5 000 spans total, dropped as soon as the
request ends healthy, 30 s TTL), degrades to plain head sampling on
overflow, and never blocks. Disable with retainTracesOnError: false.
Crashes are observed via process.uncaughtExceptionMonitor, which does
not change your process's exit behaviour; the final flush is
best-effort. Framework instrumentations are opt-in:
import { ExpressInstrumentation } from "@opentelemetry/instrumentation-express";
initAutterServer({ ..., instrumentations: [new ExpressInstrumentation()] });Version compatibility: checked for you
Some features need a minimum ingester version: operation logging needs
1.4.0, memory metrics 1.3.3, endpoint latency 1.3.1. After
initAutterServer, the SDK checks the ingester once in the background
(GET /v1/compat, unref'd, 3 s timeout). For each feature in use that the
ingester can't store, it prints one warning:
[autter-runtime] Operation logging needs ingester >= 1.4.0; yours is 1.3.4. Upgrade the ingester: docker pull ghcr.io/autter-dev/otlp-ingester:latest and restart it (ClickHouse migrations run at boot). See …/docs/COMPATIBILITY.mdThe check never throws, never delays startup or exit, and is silent when
versions match or are unknown. debug: true logs the result. Turn it off
with compatCheck: false or AUTTER_COMPAT_CHECK=0. The browser relay does
the same for browser features such as CSP violation capture. The SDK reports
itself through the OTLP resource attributes telemetry.distro.name and
telemetry.distro.version, so the Autter dashboard shows the SDK version
each service runs.
For a full report (for example in CI or after a deploy):
npx @autter/runtime-node doctor --endpoint https://ingest.example.com [--key "$AUTTER_RUNTIME_KEY"] [--json]Exit codes: 0 compatible, 1 mismatch or rejected key, 2 ingester unreachable. See docs/COMPATIBILITY.md.
Lifecycle: never lose telemetry to a forgotten shutdown()
Telemetry is batched (errors every ~2 s, healthy traces every ~5 s, metrics
every 60 s) — so exiting without flushing loses whatever is still buffered.
initAutterServer therefore installs an exit flush by default: on
beforeExit, SIGINT, and SIGTERM it force-flushes every exporter, then lets
your process die as it would have (conventional 130/143 codes). If your own
code also handles those signals, Autter only flushes alongside it and never
touches your exit path. Opt out with autoFlush: false.
Prefer explicit control? Do the same yourself and get a handle back:
import { installAutterAutoFlush } from "@autter/runtime-node";
const handle = installAutterAutoFlush(); // uses the active server's exporters
// …later: handle.flush("deploy-drain") or handle.dispose()If the process still exits with captures that were never confirmed
exported (e.g. a flush timed out), you get a one-line stderr warning — not
silent loss. While wiring things up, set debug: true (or AUTTER_DEBUG=1)
to see [autter] exported N span(s) lines on stderr as batches leave.
Note on usage rollups: requests are counted from the http.server.duration
metric (100% accurate) and additionally from sampled server spans. At the
default 1% sampling the span contribution is negligible; if you set
traceSampleRate: 1 in development, expect request counts roughly doubled.
Tail-retained error traces don't distort this: their spans carry
autter.tail_retained and the ingester keeps them out of span-fed rollups.
Privacy: secrets and PII are scrubbed before they leave the process
Error context is where secrets leak: connect ECONNREFUSED
postgres://admin:hunter2@db… in an exception message, a provider error
echoing an API key, a ?token= in a request URL, a stray
captureException(err, { "user.email": … }). Everything the SDK exports is
therefore scrubbed by default — custom attributes, exception messages and
stack traces, span status messages, LLM call attributes and errors,
structured logs, and the browser events the relay forwards. A final pass at
export time scrubs every span regardless of who created it (HTTP
instrumentation URLs, third-party instrumentations, recordException in
your own code).
- secrets inside any string are masked in place, keeping the rest of the
message/stack readable: JWTs,
Bearer/Basiccredentials,Authorization:/Cookie:/Set-Cookie:header text,scheme://user:pass@hostconnection strings (postgres, mysql, mongodb+srv, redis, …), vendor keys (sk-…,sk_live_…,ghp_…,github_pat_…,xox?-…,AIza…,AKIA…, …), PEM private keys,password=/?token=/?api_key=-style assignments, emails, and Luhn-valid card numbers; - attributes whose key looks sensitive (
password,token,secret,api_key,authorization,cookie,session,ssn,card_number, …) are masked wholesale, at any nesting depth; - non-sensitive keys, numbers (including LLM token counts) and booleans pass through untouched, so grouping, dashboards and cost tracking keep working.
The same patterns run in the browser SDK, the relay, the Python adapter,
and again in the ingester before storage (shared test vectors in
test-vectors/redaction.json).
Extend it (applies to messages, stacks and URLs too) or disable it per service:
initAutterServer({
...,
// redactAttributes: false, // opt out entirely
redactAttributes: {
additionalKeyPatterns: ["employee_id"],
additionalValuePatterns: [/^ACC-\d+$/],
},
});Libraries that must never forward PII regardless of host configuration can wrap once:
import { makeSafeCapture } from "@autter/runtime-node";
const safe = makeSafeCapture();
safe.captureException(err, { "user.email": email }); // masked before exportThe raw primitives are exported too (redactAttributes(attrs, options),
redactText(text, options)). The browser relay takes the same options as
createBrowserRelayHandler({ ..., redact }).
This closes the server-side gap to match the browser relay's payload
whitelist; it is best-effort scrubbing of obvious PII shapes, not a DLP
engine — keep secrets out of attributes in the first place.
3. LLM tracing
For the field contract, manual reporting patterns, and provider-agnostic OpenTelemetry setup, see the dedicated LLM instrumentation guide.
initAutterServer initialises LLM tracing automatically: any GenAI span —
gen_ai.* semconv attributes or the Vercel AI SDK's ai.* spans — bypasses
head sampling, so every model call is recorded with model, tokens,
latency, and a USD cost. Opt out with llmTracing: false.
Easiest: wrap the client once — works with the OpenAI, Anthropic, and Google GenAI SDKs (or anything with the same call shapes), streaming included; every call through it is traced with no per-call code:
import { instrumentLlmClient } from "@autter/runtime-node";
const openai = instrumentLlmClient(new OpenAI());
// use it exactly as before — chat, embeddings, streams are all recorded
const out = await openai.chat.completions.create({ model: "gpt-5-mini", ... });Provider is detected from the client (override with
{ provider, userId, attributes } as the second argument). For streamed
responses the span closes when the stream is consumed; OpenAI streams only
report token usage when you pass
stream_options: { include_usage: true }.
Vercel AI SDK — just turn on its telemetry, nothing else:
const { text } = await generateText({
model: openai("gpt-5-mini"),
prompt,
experimental_telemetry: { isEnabled: true, metadata: { userId: user.id } },
});Manual control (raw fetch, unusual clients) — wrap the call:
import { withLlmCall } from "@autter/runtime-node";
const res = await withLlmCall(
{ provider: "openai", model: "gpt-5-mini", userId: user.id },
async (llm) => {
const out = await openai.chat.completions.create({ ... });
llm.setUsage({
inputTokens: out.usage?.prompt_tokens,
outputTokens: out.usage?.completion_tokens,
});
return out;
},
);Errors thrown inside are rethrown after marking the span failed — failing
model calls surface both as error issues and as status: "error" LLM calls.
Costs are estimated ingest-side from a built-in price table; report exact
figures with llm.setCost(usd) (the autter.llm.cost_usd attribute).
Where wrapping is awkward (queues, callbacks, batch results), report after
the fact with
trackLlmCall({ provider, model, inputTokens, outputTokens, durationMs }).
To verify the pipe end-to-end without calling a real model:
import { emitLlmSelftestTrace } from "@autter/runtime-node";
const { traceId } = await emitLlmSelftestTrace();
// one fake "autter-selftest" call is flushed to the ingester; look it up by
// traceId in the dashboard's LLM tab (or runtime_llm_calls when self-hosting)4. Process spans (jobs, consumers, crons)
Non-HTTP work is only visible to the slow-process monitor where a span
exists — and regular traces are 1% sampled. withProcessSpan records a
span always:
import { withProcessSpan } from "@autter/runtime-node";
await withProcessSpan("invoice.rebuild", async () => {
await rebuildInvoices();
});Use stable, low-cardinality names; put ids in attributes
(withProcessSpan("email.digest", fn, { "user.id": id })).
Note on the slow-process monitor: Autter flags HTTP routes from the unsampled request metrics, so route detection works out of the box. Non-HTTP work is only visible where a span exists — relying on 1%-sampled regular traces there would undercount ~100×, which is why these spans skip head sampling.
