@omob/otel-kit
v0.10.0
Published
Configurable OpenTelemetry bootstrap for Node services: traces, metrics and logs with pluggable exporters
Maintainers
Readme
@omob/otel-kit
OpenTelemetry setup for Node services, in one function call.
Observability rests on three signals, and each answers a different question:
| Signal | What it is | The question it answers | | --- | --- | --- | | Metrics | counts and latencies, aggregated | Is something wrong? | | Traces | one request, followed span by span | Where is it wrong — which service, query or call? | | Logs | the lines your code writes | Why is it wrong, once you know where to look? |
You work down the table: a latency spike at 14:02 in the metrics leads you to the traces from that minute, and the trace_id on each log line leads you from a trace to its logs.
Tracing is the one that connects the other two, and it is what this package is mostly for. You pick where traces, metrics and logs go; it handles the SDK, the sampling, the shutdown flush and the boilerplate around spans, and stamps trace_id into your logs so every line leads back to its trace.
Install
npm install @omob/otel-kit @opentelemetry/apiThat is everything for most setups. OTLP — protobuf, JSON and gRPC — and Prometheus are already included.
Extra packages are needed for Google Cloud only, and which one depends on the route you take:
| If you use | Install |
| --- | --- |
| ExporterType.GCP for traces | @google-cloud/opentelemetry-cloud-trace-exporter |
| ExporterType.GCP for metrics | @google-cloud/opentelemetry-cloud-monitoring-exporter |
| Google Cloud over OTLP | google-auth-library — and neither of the above |
They are independent: exporting traces to Google needs the trace package only. Google is deprecating both in favour of the OTLP route, which is covered under recipes.
If you pick an exporter whose package is not installed, startup fails and names the package.
Quick start
Create src/instrumentation.ts:
import "dotenv/config";
import { ExporterType, Telemetry } from "@omob/otel-kit";
Telemetry.start({
serviceName: "my-service",
serviceVersion: process.env.APP_VERSION,
environment: process.env.NODE_ENV,
enabled: process.env.NODE_ENV !== "test",
traces: { exporter: ExporterType.CONSOLE },
// logs stay off until you add their block. Uncomment to bridge your existing pino or
// winston output, with its trace id, without changing how you log.
// logs: { exporter: ExporterType.OTLP, otlp: { url: process.env.OTEL_EXPORTER_OTLP_LOGS_ENDPOINT } },
instrumentation: { ignoreIncomingPaths: ["/health"] },
});CONSOLE needs no infrastructure — spans print to stdout, so you can confirm tracing works before you have anywhere to send it. Swap it for a real destination once you do:
traces: {
exporter: ExporterType.OTLP,
otlp: { url: process.env.OTEL_EXPORTER_OTLP_TRACES_ENDPOINT },
}Once traces go over OTLP, metrics start flowing to the same place too — see Metrics and logs for what that sends and how to stop it.
Nothing listens on port 4318 unless you run something there. If you point at a collector that is not up, every export fails with ECONNREFUSED and you see no error at all — OpenTelemetry's internal logging is off by default. The quickest real destination is Jaeger, which ingests OTLP directly:
docker run -d --name jaeger -p 16686:16686 -p 4318:4318 jaegertracing/all-in-one:1.62.0Then set OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=http://localhost:4318/v1/traces and open the UI at http://localhost:16686. Jaeger stores traces only, so also set OTEL_METRICS_EXPORTER=none, or the kit keeps trying to send it metrics it has nowhere to put.
Load it before your app:
{ "scripts": { "start": "node --require ./dist/instrumentation.js dist/server.js" } }If your app is ESM ("type": "module"), use --import instead, and keep it on the command line rather than as an import at the top of your entry file:
{ "scripts": { "start": "node --import ./dist/instrumentation.js dist/server.js" } }ESM links every module in the graph before any of them runs, so an import "./instrumentation.js" inside server.js starts telemetry after Fastify, ioredis or kafkajs have already loaded — too late to patch them. --import runs first. Node 18.19 or later is needed for ESM instrumentation; on older runtimes only CommonJS requires are patched.
That's it. HTTP, database and framework calls are traced automatically.
Every trace is kept until you say otherwise. Keep it that way while you are setting things up: sampling is a production concern, and turning it down before you have seen a single trace is the most common reason nothing appears in a backend. When you're ready, set OTEL_TRACES_SAMPLER_ARG=0.1 to keep 10% — the kit reads the standard variable itself and ignores a value that isn't between 0 and 1, with a warning, rather than drop every trace. A traces.sampleRatio in code wins over it.
On Fastify, install @fastify/otel and turn it on. It ships disabled, along with fs:
instrumentation: { enable: [InstrumentationName.FASTIFY], ignoreIncomingPaths: ["/health"] }Without it, every request is one bare GET span with no http.route, so nothing groups by route. Express, Koa, Hapi, NestJS, Mongo, Postgres, Redis, Kafka and outbound HTTP need no such step — they are on by default. Fastify is different because the OpenTelemetry-owned instrumentation was deprecated in favour of the Fastify team's own @fastify/otel, which this package loads when you enable it (npm i @fastify/otel).
Why --require
Instrumentation can only patch libraries loaded after it starts. --require guarantees that.
Importing it at the top of your entry file also works, as long as nothing you want traced is imported above it. One reordered import and tracing silently stops — hence the flag.
What every signal carries
Every span, metric and log is stamped with who sent it, so a backend can tell services, versions and replicas apart:
| Attribute | Where it comes from |
| --- | --- |
| service.name | serviceName |
| service.version | serviceVersion, or, if you leave it out, the version in the package.json npm or pnpm ran your start script from — so each deploy shows up as a new version |
| service.instance.id | a random id per process, so replicas are counted separately; set your own through resourceAttributes or OTEL_RESOURCE_ATTRIBUTES, such as the pod name |
| deployment.environment.name | environment |
| ritele.trace.sample_probability | the chance this service keeps a trace it starts, when it exports traces; see below |
| telemetry.sdk.* | the OpenTelemetry SDK's name, language and version |
| host.name, host.id | detected at startup. On a personal machine — macOS, Windows, or Linux with a desktop session or under WSL — both are replaced by stable pseudonyms such as host-3fa9c1d2e4b7, because the name is usually its owner's (Ada-MacBook-Pro) and the id is a hardware fingerprint. Linux servers, containers and pods send them as they are. See hostName below |
| host.arch, process.pid, process.runtime.* | detected at startup |
| container.id | detected at startup under Docker and on hosts with cgroup v1. Most current Kubernetes clusters (containerd with cgroup v2, as on EKS, GKE and AKS) don't expose it to the process, so it is left out there |
| k8s.pod.name, k8s.namespace.name, k8s.container.name, … | not detected — pass them in through OTEL_RESOURCE_ATTRIBUTES, as the Kubernetes recipe shows. Pod, namespace and container name are what the cluster's own metrics are labelled with, so they are the attributes to set |
Ids in URLs are masked. An account number, email, UUID, token or document number in a request path or query string becomes * in url.path, url.query and url.full before a span leaves the process: /v1/customers/[email protected]&page=2 is sent as /v1/customers/*?email=*&page=2. http.route keeps the route template (/v1/customers/:id), so grouping by route loses nothing. The rules were checked against every path segment of a large production codebase, so route names such as process-multi-payment-wallet-credit-retry or confirm-otp-v2 stay readable. They are rules, not a guarantee: a short technical name with digits (sha256, base64, top-10) is masked too, and an id made only of letters and shorter than 24 characters is not. Query values such as a person's name have no shape to match, so if yours carry them, set traces.redactQuery: QueryRedaction.DROP to send no query strings at all. Set traces.redactPathSegments: false if you need raw URLs. Masking runs as every span ends — failed and aborted requests included, unlike a hook that only sees responses — and applies to url.* values whatever set them, your own instrumentation hooks included, so a hook that blanks the query keeps it blank. DROP applies even with redactPathSegments: false.
Personal details stay on the machine. The kit leaves out your command line, the paths to your script and to the node binary, and your user name, and on a personal machine it swaps the host name and id for pseudonyms. Flags often carry secrets, and paths such as /Users/<you>/.nvm/… and host names such as Ada-MacBook-Pro name the user. A pseudonym is the same on every run, so a backend still tells your machines apart and still has a label to show. The host name's pseudonym is keyed by the machine id, which never leaves the machine, so it can't be reversed by guessing likely names; a personal machine with no machine id sends no host name.
You stay in control. hostName decides the host name and id: HostNameMode.AUTO (the default, as above — which also treats macOS and Windows servers as personal, and Linux only with a graphical session or under WSL; a terminal-only Linux laptop needs HASH), KEEP to send them as they are, HASH to always use pseudonyms, or HIDE to send no host name at all. OTEL_NODE_RESOURCE_DETECTORS still chooses which detectors run (env, host, os, process, serviceinstance, container, all or none), and the same protections apply to whichever you pick. The kit's all also includes container.
In a monorepo, check service.version. A start script run from the repo root reports the root package.json's version (often 0.0.0) for every service. Set serviceVersion yourself there.
Your own spans
withSpan runs your function inside a span. It starts the span, makes it the parent of anything that happens inside, ends it when your function settles, and records the error if one is thrown.
Add one when you want a step to show up as its own line in the trace: a slow query, an external API call, a step you suspect. Skip it for cheap in-memory work — a span costs more than the code it measures.
import { withSpan } from "@omob/otel-kit";
async function login({ email, password }) {
return withSpan("login", { attributes: { "auth.method": "password" } }, async (span) => {
const user = await findUser(email);
span.setAttribute("user.id", user.id);
return withSpan("token.generate", () => generateToken(user));
});
}Throw anywhere inside and the span is marked failed, the exception is recorded, and the error still propagates to your caller unchanged. Spans always end, on success or failure.
Two shorthands help when a backend is drawing a dependency graph from your spans: peer names the remote side of a call that has no instrumentation of its own, and component names the logical part of your service the work belongs to.
import { SpanKind } from "@opentelemetry/api";
declare const paystack: { charge(order: { id: string; amount: number }): Promise<unknown> };
async function charge(order: { id: string; amount: number }) {
return withSpan("charge.card", { kind: SpanKind.CLIENT, peer: "paystack", component: "billing" }, () =>
paystack.charge(order)
);
}Describing your architecture
If you run something that builds a system diagram from traces, tell it what this service is. None of this changes what is traced; it adds resource attributes and a second, independent sampling decision.
Telemetry.start({
serviceName: "wallet-service",
traces: { exporter: ExporterType.OTLP, sampleRatio: 0.01 },
architecture: {
component: { type: ArchitectureComponentType.SERVICE, layer: "core", domain: "payments", owner: "team-wallet" },
intendedDependencies: ["postgresql:ledger", "kafka:transfers", "paystack"],
concurrency: { http: 200, pgPool: 20 },
docTraceRatio: 0.02,
peers: { "api.paystack.co": "paystack" },
},
});docTraceRatio is the useful one. Sampling 1% of traffic keeps costs down, but a rarely-used dependency can go unseen for days. Documentation traces are a separate 2%, taken from the end of the range your sample ratio never reaches, always recorded, and marked in tracestate so every service downstream records them too. A backend can keep those at 100% and drop the rest, and the map stays complete.
Each exporting service also carries ritele.trace.sample_probability: how likely it is to keep a trace it starts. That is sampleRatio + docTraceRatio, capped at 1 — the two ranges never overlap, so they simply add. With the settings above it is 0.03, so a backend can multiply a count of this service's root spans by about 33 to estimate the real number. It says nothing about traces that arrive from a caller: those are kept or dropped by the caller's decision. The attribute is left out when you pass your own traces.sampler, since the kit can't know its rate.
The attribute names live under ritele.*, the namespace of Ritele; any other backend ignores them.
Not every failure is a fault. A wrong password is an expected outcome, and marking it as a span error means your error rate tracks how often users mistype. Pass isError to say which throws actually count:
withSpan("login", { isError: (e) => !(e instanceof AppError) || e.statusCode >= 500 }, handler);Never put emails, tokens or passwords in attributes — spans are stored unredacted.
Two more helpers:
import { currentTraceId, getTracer } from "@omob/otel-kit";
currentTraceId(); // trace id of the active span, or undefined
getTracer("auth-module"); // pass as `tracer` in withSpan options to name the scopecurrentTraceId() is worth putting in your error handler, so support can jump from an error response straight to the trace:
fastify.setErrorHandler((err, request, reply) =>
reply.status(err.statusCode ?? 500).send({ message: err.message, traceId: currentTraceId() })
);Metrics and logs
Neither needs code beyond the config.
Metrics
If your traces go over OTLP, metrics are already on. You don't add anything: every 30 seconds the kit sends metrics to the same collector as your traces. Most collectors, and backends such as Ritele, accept both.
They carry the headers you set in code (traces.otlp.headers) or in OTEL_EXPORTER_OTLP_HEADERS. If your auth lives only in OTEL_EXPORTER_OTLP_TRACES_HEADERS, it is not copied: set OTEL_EXPORTER_OTLP_METRICS_HEADERS as well, or the metrics are sent without it and rejected.
Where exactly they go:
| Your setup | Metrics are sent to |
| --- | --- |
| traces.otlp.url ends in /v1/traces | the same URL, ending in /v1/metrics instead |
| the traces URL comes from OTEL_EXPORTER_OTLP_TRACES_ENDPOINT | the same, from that variable |
| OTEL_EXPORTER_OTLP_METRICS_ENDPOINT is set | that endpoint, with only the headers you gave it in OTEL_EXPORTER_OTLP_METRICS_HEADERS or OTEL_EXPORTER_OTLP_HEADERS — never your traces headers, since it may be another vendor |
| gRPC | the same endpoint as traces |
| no URL in code or in either variable above | wherever your traces go: OTEL_EXPORTER_OTLP_ENDPOINT if set, otherwise the local default (localhost:4318, or 4317 for gRPC) |
| a traces URL with any other path | nowhere — the kit can't guess, so metrics stay off |
To turn them off, set OTEL_METRICS_EXPORTER=none, or pass metrics: { exporter: ExporterType.NONE }. Do this if your backend only takes traces (Jaeger, for example) or charges per metric series. The variable wins over anything in code, including a metrics block, so it works as an off switch during an incident; the kit prints a warning at startup when it overrides a block, so a leftover setting doesn't go unnoticed. Only none is read; other values such as console are ignored.
To change an option, such as the interval or cpuUsage, write a metrics block with exporter: ExporterType.OTLP and no URL. It keeps the collector and headers worked out above. If there is none to keep — traces don't go over OTLP, or their URL has another path — the block uses OTEL_EXPORTER_OTLP_METRICS_ENDPOINT or OTEL_EXPORTER_OTLP_ENDPOINT if you set one. Otherwise the kit sends no metrics and says so at startup, rather than quietly retrying a local collector that isn't there. To send them somewhere else, give the block its own URL:
metrics: {
exporter: ExporterType.OTLP,
otlp: { url: process.env.OTEL_EXPORTER_OTLP_METRICS_ENDPOINT },
exportIntervalMillis: 30_000,
}If your traces don't go over OTLP, metrics stay off until you add that block.
What you get, with no instrumentation of your own:
| Metric | What it tells you |
| --- | --- |
| http.server.request.duration, http.client.request.duration | Request latency in and out, in seconds, by route and status |
| nodejs.eventloop.delay.p50 / p90 / p99, nodejs.eventloop.utilization | How busy the event loop is — usually the first thing to saturate on a Node service |
| v8js.memory.heap.used, v8js.memory.heap.space.* | How much of the heap is in use, per heap space |
| ritele.v8js.memory.heap.limit | The heap's ceiling, in bytes — the size at which Node runs out of memory. Chart v8js.memory.heap.used against it to see how close a service is. OpenTelemetry has no standard name for this (its old v8js.memory.heap.limit meant something else and is retired), so it lives under ritele.* |
| messaging.client.sent.messages, messaging.client.consumed.messages, messaging.process.duration | Kafka throughput and handler time, if you use kafkajs |
The event loop and heap metrics, the heap ceiling included, stay on even when you narrow things down with instrumentation.only. Set runtimeMetrics: false to drop them.
Dashboards on
http.server.duration? That is the old name, in milliseconds, from earlier versions of the HTTP instrumentation. The one bundled here reports onlyhttp.server.request.duration, in seconds, andOTEL_SEMCONV_STABILITY_OPT_INno longer switches it back. Point those dashboards and alerts at the new name.
For a pull-based setup, swap the exporter and Prometheus scrapes you instead:
metrics: { exporter: ExporterType.PROMETHEUS, prometheus: { port: 9464 } }Logs
Logs are off until you add a logs block:
logs: { exporter: ExporterType.OTLP, otlp: { url: process.env.OTEL_EXPORTER_OTLP_LOGS_ENDPOINT } }You do not change how you log. If you use pino, winston or bunyan, the log instrumentation bridges what you already write into OpenTelemetry, carrying the trace_id that ties each line to its span — so a trace links straight to the logs from that request.
Two things to weigh before turning logs on. Your log volume goes to two places, so you pay to store it twice unless you drop stdout collection. And any gap in your redaction now reaches a second system: check what your logger emits — response bodies and auth headers are the usual leaks — before pointing it at a backend.
Connection pools
A slow query and a query stuck waiting for a free connection look the same from outside. The difference only shows inside your process, where the pool knows its limit (max: 20) and how many callers are queued. Register the pool and the kit reports it:
import { observeConnectionPool } from "@omob/otel-kit";
const pool = new Pool({ max: 20 });
observeConnectionPool({
name: "biller",
system: "postgresql",
read: () => ({
max: pool.options.max,
used: pool.totalCount - pool.idleCount,
idle: pool.idleCount,
pending: pool.waitingCount,
}),
});Call it after Telemetry.start(). A pool registered earlier gets a do-nothing meter and reports nothing, for good.
You get:
| Metric | What it tells you |
| --- | --- |
| db.client.connection.max | The pool's limit |
| db.client.connection.count, split into used and idle | How much of it is in use |
| db.client.connection.pending_requests | Callers waiting for a connection |
| db.client.connection.wait_time | How long they waited, if you call the returned recordWait(millis) when you acquire one |
A pool at its limit with a queue behind it means the bottleneck is your pool size, not the database. It works with any pool — Postgres, MySQL, Mongo, Redis — because you supply the read function.
Several pools, one database? Register each under the same name and their readings add into one series, which is what the database sees from your service. A primary and a replica that point at the same host when no replica is configured are the usual case. Stopping one leaves the others reporting. The knex recipe shows this for knex.
Using pg? The pg instrumentation also reports these metrics for pg-pool, under the same names. Its numbers are only right while you have a single pool: with two or more, its counts drift and can go negative. Register your pools here anyway, and have your backend read the @omob/otel-kit scope. Ritele prefers these measured numbers over architecture.concurrency.pgPool, which it only uses for a pool nothing measures.
CPU capacity
To predict when a service runs out of CPU, a backend needs how much CPU each process uses and how much each replica is allowed. Turn on the first with metrics.cpuUsage and state the second with architecture.cpuLimit:
Telemetry.start({
serviceName: "wallet-service",
traces: { exporter: ExporterType.OTLP, otlp: { url: process.env.OTEL_EXPORTER_OTLP_TRACES_ENDPOINT } },
metrics: { exporter: ExporterType.OTLP, cpuUsage: true },
architecture: { cpuLimit: 0.5 },
});That emits process.cpu.time, split by cpu.mode, and the resource attribute ritele.cpu.limit. Replicas are counted from service.instance.id, which the kit sets to a random id per process unless you supply one. To observe CPU without the config flag, call observeCpuUsage() after Telemetry.start(); it returns { stop }. See Recipes for reading the limit from Kubernetes.
More
- Configuration — every option, shutdown behaviour, and what happens when a config is rejected
- Recipes — Jaeger, Google Cloud, Prometheus, gRPC collectors, per-instrumentation options
- Troubleshooting — no traces appearing, wrong service name, broken propagation
- Concepts — traces, spans, sampling and propagation, if OpenTelemetry is new to you
Licence
MIT
