@genn-inc/cluebase-backend-sdk
v0.0.2
Published
Cluebase Backend SDK (Node.js) — OTel-based observation-source event capture for Express, NestJS, MCP, LangChain, Mastra, and AI providers
Readme
@genn-inc/cluebase-backend-sdk
Cluebase Node.js / Express backend SDK. Events flow through a
NodeTracerProvider + consent-aware span processor +
ObservationSourceEventExporter to
/api/v1/ingest/backend.
Preload the instrumentation (required)
Start the server with the SDK's preload entry:
node --import @genn-inc/cluebase-backend-sdk/register dist/main.jsAdd the same flag to every start path — each start / start:* / dev script,
plus the Dockerfile CMD, PM2, or systemd unit if the server is launched there.
NODE_OPTIONS=--import @genn-inc/cluebase-backend-sdk/register works when editing the
command is not practical.
Without it the SDK still starts and reports no error, but nothing is recorded
about what the backend did: requests, database reads and writes, branch
decisions, and outbound dependency calls all stay empty. Node instrumentation
patches a module while that module is loading, so it cannot reach a module that
is already loaded — and express and the database client are loaded before
cluebase.init(...) runs. cluebaseExpressMiddleware annotates the request span that
this instrumentation creates; it never creates one itself.
The preload installs @opentelemetry/instrumentation-http, -pg, and
-undici, which ship with this package. The bundled
@opentelemetry/auto-instrumentations-node set is deliberately not used: it
also records file-system, DNS, and socket activity, which buries the
observations that matter. Configuration stays with cluebase.init(...) — the
preload installs instrumentation only.
Database coverage is PostgreSQL only today. Requests, branch decisions, and
outbound dependency calls are recorded for any stack, but reads and writes are
recorded only through pg. On MySQL, MongoDB, Redis, or Prisma the
data-operation side stays empty. Add the matching OTel instrumentation to your
own bootstrap alongside the preload if you need it:
import { registerInstrumentations } from "@opentelemetry/instrumentation";
import { MySQL2Instrumentation } from "@opentelemetry/instrumentation-mysql2";
registerInstrumentations({ instrumentations: [new MySQL2Instrumentation()] });Place that in a module you preload the same way, so it also runs before the database client is loaded.
Minimal integration
import cluebase from "@genn-inc/cluebase-backend-sdk";
cluebase.init({
projectKey: process.env.CLUEBASE_PROJECT_KEY!,
apiKey: process.env.CLUEBASE_API_KEY!,
// Literal service label from Cluebase setup output. Do not expose this to the browser.
serviceKey: "backend-api",
});
cluebase.identify(user.id, { name: user.name, email: user.email });
cluebase.group("organization", organization.id, {
name: organization.name,
customer_defined_segment: organization.segment,
});
cluebase.track("order_placed", { product_name: "shirt", amount: 100 });
// Call on logout/session reset success paths.
cluebase.reset();
await cluebase.flush();identify and group keep arbitrary privacy-safe customer attributes as
nested occurrence-time subject_traits and organization_traits. Profile-safe
identity values (display_name, sanitized avatar_url, contact hash/domain)
remain separate from those generic traits. Request-scoped contextTraits
supplied through the advanced request-context API are captured as
context_traits. These maps are snapshotted when a permitted span starts; they
are not flattened or JSON-encoded into OTel attributes.
Do not pass environment; Cluebase derives dev/prod from the projectKey prefix.
If the customer product calls its company concept account, workspace, tenant, or
team, pass that stable company id to cluebase.group("organization", ...) for MVP.
endpoint is optional and defaults to https://api.cluebase.io. Set
CLUEBASE_API_BASE_URL only when the SDK must send to another Cluebase API base
URL, such as a local environment.
cluebase.track custom event names must match
^[A-Za-z0-9_.:-]{1,128}$. Invalid names are ignored and do not emit a
source-signal custom intent.
Privacy and PII handling
1. Hard-deny: PII / secrets are stripped before transport
The SDK strips a built-in set of property keys before the value is
set as an OTel span attribute (i.e. before it leaves the customer
process). Caller-supplied deniedKeys are added on top of the
hardcoded ALWAYS_DENIED_KEYS — they cannot remove a default-denied
key.
Hard-deny categories (case- and separator-insensitive — userEmail,
user-email, USER_EMAIL, email_address all match):
- Auth credentials:
authorization,cookie,set-cookie,password,passwd,secret,token,access_token,refresh_token,session,session_token,api_key,apikey,private_key - PII categories:
email,phone,credit_card,ssn
This list is a strict superset of the server-side ingest hard-deny
(@cluebase/shared INGEST_HARD_DENY_KEYS), enforced by a CI parity test
in test/sanitize.test.ts.
Nested objects in cluebase.track properties are recursively sanitized —
cluebase.track("x", { user: { email: "[email protected]", name: "Alice" } }) emits
the name field but never transmits email.
2. Analysis projection and later registration
Properties that pass the hard-deny gate are projected by default unless a
cardinality, free-text, or complex-value guard keeps them out of the analysis
index. Guarded non-private values remain available in
analysisProperties.raw_only so they can be registered and promoted later.
The map stores each original value as JSON so its type can be restored.
Hard-denied values are dropped and never stored in raw_only.
After a privacy review, use the Cluebase admin UI or project service allowlist API to promote a guarded non-private property into the analysis projection.
Architecture boundary
The Node.js SDK contains a shared instrumentation core and framework/provider adapters. Adapters register native hooks and extract framework/provider facts; the shared SDK core owns span lifecycle, correlation context, privacy handling, common attributes, failure status, and transport. The exporter serializes raw source signals, and the Cluebase backend ingest core performs canonical classification, normalization, privacy projection, and raw-ingest shaping.
Node.js framework / library hook
-> native facts passed to shared SDK instrumentation core
-> OTel source signal
-> ObservationSourceEventExporter
-> /api/v1/ingest/backend
-> apps/api/src/modules/ingest canonical coreRules:
- Keep adapters limited to native hook registration, native fact extraction, and calls into the shared SDK core.
- Keep run/step tracking, common span lifecycle, context projection, privacy policy, failure classification, and delivery outside framework/provider adapters.
- Keep canonical event classification in the backend ingest core so every SDK language uses the same event rules.
- Preserve
interaction_id,request_span_id,request_id, andtrace_idwhenever they are available through the shared SDK core. - Run
pnpm sdk:architecture:checkwhen changing SDK source. CI runs this gate across every SDK registered in the descriptor.
The stable boundary contract is
architecture-gates.md.
Advanced OTel integrations must install the exporter's consent-aware processor:
const exporter = new ObservationSourceEventExporter(settings);
const provider = new NodeTracerProvider({
spanProcessors: [exporter.createSpanProcessor()],
});A standard BatchSpanProcessor cannot prove consent at span start. After a
consent generation changes, direct standard-processor export therefore fails
closed until the exporter is reconstructed.
Build / test
pnpm install
pnpm build
pnpm test # unit + parity tests
pnpm test:cov # with coverage