@geonosis/observability
v3.1.0
Published
Ports whose answer can say failed: an ErrorSink and a Tracer returning Result, one attribute vocabulary (tenant as a group id, never a person), a memory sink, envelope-id dedupe, a PostHog adapter, a tiered health registry, and a probe that plants an erro
Maintainers
Readme
@geonosis/observability
Through the front door: geonosis observability — the metapackage pins this and every other kit
tool at ONE version, and passes the exit code through unchanged.
Ports whose answer can say failed, a vocabulary that refuses a person as a tenant, and a probe that plants an error and goes and looks for it.
Extracted from a Workers + D1 app's PostHog sink and health endpoint, a Medusa storefront's provider seam, and Midday's
tiered health registry — after an afternoon in which four instruments reported green while measuring
nothing (docs/rules-backlog.md #42). Zero runtime dependencies. Node ≥ 22 and workerd: every flush
is explicit or returned as a promise the host can waitUntil, and no timer is ever left running.
pnpm add @geonosis/observability better-result
pnpm add posthog-node # only if you use the /posthog adapter| Import | What is behind it |
|---|---|
| @geonosis/observability | the ports, the vocabulary, createMemorySink, createSwallowingSink, createMemoryTracer, withDedupe, withLastEvent, proveObservability |
| @geonosis/observability/health | createHealthRegistry — tiers, TTL, timeouts, ready() vs dependencies() |
| @geonosis/observability/posthog | createPostHogSink — the vendor is injected, never imported |
| geonosis-observability prove | the CLI. Four verdicts, exit 0 or 2 |
1. The ports
type ErrorSink = {
capture: (error: Error, attributes: Attributes) => Promise<Result<CaptureReceipt, SinkFailure>>
flush: () => Promise<Result<void, SinkFailure>>
kind: string
}
type CaptureReceipt = { deduplicated: boolean; id: string }better-result is a required peer: Result is in every return type, so you cannot call this
package without it, and a second copy in a dependencies field would put two Ok classes in one
tree.
Why not Promise<number>. a Workers + D1 app's seam answers with a count: events.length from the real
sink, 0 from the one with no key, a rejected promise from a dead transport — handed to
waitUntil, where nothing reads it. A count cannot say failed. That is the shape no-void-port
refuses, wearing a number instead of void.
Why the receipt carries an id. React strips an error's message on the way to a browser; a deploy renames every stack frame. The id the sink filed the report under survives both, which is why it is in the answer rather than in a log line.
const receipt = await sink.capture(error, { envelopeId, service: 'api', tenant: orgId })
if (receipt.isErr()) log.warn('the report did not land', receipt.error.sink, receipt.error.message)
else response.headers.set('x-error-id', receipt.value.id)The tracer
tracer.span('queue.handle-batch', { service: 'api', tenant: orgId }, async () => doTheWork())Returns whatever the body returned and rethrows whatever it threw. It is a recorder, not a gate: a tracer that swallowed a throw would make a failed request answer 200. Its own failures are its own business and never the caller's.
Span names are <area>.<action>, lowercase, at least two dot-separated segments, dashes inside a
segment: queue.handle-batch, db.cell.query, http.request. isSpanName(name) decides;
SPAN_NAME_CONVENTION states it. Never enforced by a throw — a bad span name must not become a 500
— so assert it in your own tests if you want it held.
2. The vocabulary
| Field | Meaning |
|---|---|
| service | required. A report nobody can route to a service is a report nobody reads. |
| tenant | Always a group id, never a person. Refused if it contains @. |
| envelopeId | The id this occurrence already has upstream. The dedupe key, and the id it is filed under. |
| runtime | node, workerd, browser, a container id. An open string. |
| release | Which build. |
| actor | The one field that MAY carry an identity. |
| extra | Anything else. May not shadow a field above. |
release and runtime are here because neither consumer attaches either, and a production error
that cannot be pinned to a deploy is the first of the four lying instruments.
tenant refusing an @ is the one rule that is not a shape check. A group id is the join key every
dashboard fans out on, and the one field an erasure request cannot reach once an address is in it —
and it is the mistake both repos would make first, because distinctId legitimately holds a person
and the two fields sit one line apart.
import { parseAttributes, attributesSchema } from '@geonosis/observability'
import { z } from 'zod'
parseAttributes(value) // → Result<Attributes, InvalidAttributes>, no zod anywhere
attributesSchema(z) // → the same rules, built over YOUR zodTwo doors, one set of rules. Both consumers run different zod versions (4.2.0, 4.4.3 and 4.5.4 measured across the two trees), so this package imports none.
3. The sinks
const sink = createMemorySink() // a real implementation: what tests and the probe read
const sink = createSwallowingSink() // answers ok, keeps nothing — shipped so --prove can catch itcreateSwallowingSink is a Workers + D1 app's silentAnalyticsSink with the console.info removed. It is
here on purpose: it is what both consumers actually run whenever no key is set, and a probe that
could not tell it from a real sink would prove nothing.
Dedupe
withDedupe(sink) // an in-memory Map
withDedupe(sink, { seen, remember }) // your store: a KV namespace, an audit tableAt-least-once in, at-most-once counted. A repeat of a known envelopeId answers with the id the
first capture was filed under and deduplicated: true, and the inner sink is not called. An
envelope whose capture failed is not remembered — buying at-most-once by losing at-least-once
would make one transport error permanent silence for that envelope.
An attributes bag with no envelopeId always passes through: a Workers + D1 app mints a fresh uuid for a
request-door failure precisely because it has none, and dropping the second of two would lose a
real second failure.
The last-event file
withLastEvent(sink, (record) => writeFile('.geonosis/last-event.json', JSON.stringify(record)))Writes { at, id, sink, release? } after a capture the sink accepted — nothing for a refusal,
nothing for a deduplicated repeat. This is the file geonosis-doctor's observability check reads.
The writer is injected because a file is a Node fact; on workerd the same record goes to KV.
4. The PostHog adapter
import { PostHog } from 'posthog-node'
import { createPostHogSink } from '@geonosis/observability/posthog'
const sink = createPostHogSink({
newClient: (batchSize) =>
new PostHog(env.POSTHOG_API_KEY, {
host: env.POSTHOG_HOST,
flushAt: batchSize,
flushInterval: 0, // no long-lived process to run a timer in
}),
groupName: 'organization', // default
identityOf: (a) => a.actor ?? `org:${a.tenant}`, // a Workers + D1 app's
})
// in a worker
ctx.waitUntil(sink.flushable().promise)posthog-node is never imported here — you pass the constructor. That keeps this package at zero
runtime dependencies, keeps your key and host out of it, and makes the whole adapter drivable
against a stub server with no network.
flushable() returns { pending, promise }. Read pending before awaiting to skip the call when
there is nothing to send; the promise resolves, never rejects, so a host handing it to
waitUntil is not handed a rejection it never asked for. Read the Result if you want to know.
Two vendor behaviours this is built around, both measured against [email protected]:
- A
uuidthat is not a valid UUID is silently replaced. a Workers + D1 app's ids arecrypto.randomUUID(), so its dedupe works; ids likeevt_7, a ULID, a nanoid or a Medusaorder_01H…would be dedupe keys that never arrive. The adapter passes a UUID through untouched and derives a stable UUIDv5 from anything else — same envelope, same uuid, every process, verified against CPython'suuid5— and puts the original on the wire asproperties.envelopeId. shutdown()resolves over a batch the server refused. A 500 and a dead port both leave the await resolving normally; the failure is reported only through anon('error')listener. The adapter subscribes and turns what it hears into theSinkFailure.
At the composition root
The sink is built ONCE, where the app is wired, and handed to everything else as the port. A repo
that installs this package and composes nothing sends production errors nowhere while every gate
stays green — the reference consumer did exactly that, which is why geonosis-doctor --only drift
now WARNs when no source file in the tree imports this package at all.
// apps/api/src/observability.ts — the composition root, imported by the entry point and nothing else
import { PostHog } from 'posthog-node'
import { withDedupe, withLastEvent } from '@geonosis/observability'
import { createPostHogSink } from '@geonosis/observability/posthog'
export const sink = withLastEvent(
withDedupe(
createPostHogSink({
newClient: (batchSize) =>
new PostHog(env.POSTHOG_API_KEY, { flushAt: batchSize, flushInterval: 0, host: env.POSTHOG_HOST }),
}),
),
// The writer is yours: a file on Node, a KV put on workerd. This package touches no fs.
async (event) => writeFile('.geonosis/last-event.json', JSON.stringify(event)),
)// geonosis.json — what the doctor reads back
"observability": {
"sink": "posthog",
"endpoint": "https://eu.i.posthog.com",
"lastEventFile": ".geonosis/last-event.json",
"maxAgeSeconds": 3600
}The entry point imports sink and nothing else constructs one: two sinks in one process are two
dedupe caches, two flush schedules, and one of them writing the last-event file the doctor reads.
The browser half is yours. posthog-js is a bundler and framework fact — a key-gated, memoised
dynamic import('posthog-js'), identify(userId) on sign-in, group('organization', orgId) so the
browser session and the server events join on the same key. Use the same group name you passed here
and the two halves meet; that is all this package can usefully say about it.
5. --prove
geonosis-observability prove --sink memory
geonosis-observability prove --sink posthog-stub
geonosis-observability prove --sink posthog --verify-command 'my-posthog-query "$GEONOSIS_TRACE_ID"'Plants an error under a trace id nothing else in the world has, then goes and LOOKS.
| Verdict | Meaning | Exit |
|---|---|---|
| PROVEN observability: the planted error reached <sink> as <trace id> | found | 0 |
| CANNOT FAIL: the sink swallowed it | answered ok, kept nothing | 2 |
| MISREAD: … | kept something, but not the plant | 2 |
| CANNOT MEASURE: … | nothing configured, a refused plant, or a failed flush | 2 |
--sink posthog refuses without --verify-command. A key proves you can send; sending is not
arriving, and "the SDK did not complain" is exactly the evidence that was found to be worthless.
Nothing here queries anybody's project — your command runs with GEONOSIS_TRACE_ID in its
environment and its exit code is the answer.
Put your own sink behind the same four verdicts with the exported proveObservability:
const proof = await proveObservability({ probe: { kind, sink, seen: (id) => lookItUp(id) } })
process.exit(exitCodeOf(proof))6. The health contract
import { createHealthRegistry } from '@geonosis/observability/health'
const health = createHealthRegistry({
dependencies: [
{ name: 'cell', tier: 1, check: () => pingCell(db), timeoutMs: 2000, ttlMs: 5000 },
{ name: 'controlPlane', tier: 1, check: () => pingControl(control), timeoutMs: 2000 },
{ name: 'search', tier: 4, check: () => pingSearch(), timeoutMs: 1000, ttlMs: 30_000 },
],
})ready() runs tier 1 only and answers { ok, status: 200 | 503, degraded, dependencies }.
dependencies() runs every tier; its status still follows tier 1, so a degraded search index does
not take a pod out. A registry with no tier-1 dependency answers 503 — a 200 from a registry that
verified nothing is the /health over an unreachable database, built out of an empty array.
timeoutMs (5s default) makes a hang a failure with timed out after Nms, and the timer is cleared
on the path where the check answered first. ttlMs caches a good answer and marks it
cached: true; a failure is never cached, because a service reported dead for a minute after it
recovered is worse than one nobody predicted.
Fixture — Hono
app.get('/health/ready', async (c) => {
const report = await health.ready()
return c.json(report, report.status)
})
app.get('/health/dependencies', async (c) => {
const report = await health.dependencies()
return c.json(report, report.status)
})Fixture — Node http
createServer(async (request, response) => {
const report =
request.url === '/health/ready' ? await health.ready() : await health.dependencies()
response.writeHead(report.status, { 'content-type': 'application/json' })
response.end(JSON.stringify(report))
})Both are examples, not code this package ships: a health registry that depended on one server could not be used by the other.
7. The doctor line
geonosis-doctor gains an observability check when geonosis.json has a block:
{
"observability": {
"sink": "posthog",
"endpoint": "https://eu.i.posthog.com",
"lastEventFile": ".geonosis/last-event.json",
"maxAgeSeconds": 3600,
"probe": "geonosis-observability prove --sink posthog-stub"
}
}It asks whether a sink is configured (a memory/noop/swallowing one is a WARN), whether the
endpoint answers a HEAD, and whether an event arrived inside maxAgeSeconds. It reads the FILE
withLastEvent wrote and imports nothing of this package — a check that needed the library it
checks cannot run in the tree where the library is missing, which is the first case it exists to
find. It names your probe command; it never runs it.
Migrating a hand-rolled mirror
A repo's own PostHog mirror module becomes a createPostHogSink instance plus withDedupe.
| Today | With this package |
|---|---|
| AnalyticsSink { capture: (events) => Promise<number> } | ErrorSink returning Result<CaptureReceipt, SinkFailure> |
| distinctId: message.actor ?? \org:${orgId}`|identityOf: (a) => a.actor ?? `org:${a.tenant}`— theorg:prefix stays yours |
|groups: { organization: orgId }|tenant: orgId+groupName: 'organization'(the default) |
|uuid: message.id|envelopeId: message.id— already a UUID, so it goes on the wire unchanged |
|new PostHog(key, { flushAt: events.length, flushInterval: 0 })per batch | yournewClient(batchSize), verbatim |
| context.waitUntil(analytics.capture(mirrored))|for (…) await sink.capture(…), then context.waitUntil(sink.flushable().promise)|
|await client.shutdown()thenreturn events.length| theResult—shutdown()resolving over a refused batch is why |
|silentAnalyticsSinkwhen no key |createSwallowingSink(), and --provesaysCANNOT FAILover it |
|handleQueue({ seen: (id) => hasAuditEntryForEvent(db, { eventId: id }) })| keep it — and addwithDedupe(sink, { seen, remember })for the counting |
|FailureDoor = 'cron' | 'mail' | 'queue' | 'request' | 'workflow'|extra: { door }— the five values are a worker's, not a vocabulary's |
|failureEventFormintingcrypto.randomUUID()| leaveenvelopeIdoff; the sink mints one |
|healthStatus(version, probes)|createHealthRegistry, with cellandcontrolPlaneat tier 1 and atimeoutMseach |
| one/health|/health/readyfromready(), /health/dependenciesfromdependencies()|
| nothing |releaseandruntimeon every capture, and aprobeline ingeonosis.json` |
apps/web/lib/analytics/browser.ts does not move. It is posthog-js, a bundler-visible env read and
a dynamic import — all three are the app's. Keep group('organization', orgId) matching groupName.
Migrating a Medusa storefront
A Medusa storefront has no error sink and no health endpoint at all. Its migration is adoption, not
replacement, and its own best-shaped seam is the template:
packages/medusa-plugins/customer-service/src/lib/llm-provider.ts resolves three ways
(anthropic | claude-cli | dev) and its no-key path is a real deterministic implementation, not
a null object.
- Put
createMemorySink()behind that same shape — a resolver that returns the memory sink with no key and a real adapter with one. A Medusa storefront's no-key path is then readable, whichsilentAnalyticsSinknever was. serviceis the plugin or the app;tenantis nothing until it has one — leave it off rather than putting a customer in it.- Medusa entity ids (
order_01H…) are not UUIDs. Pass them asenvelopeIdanyway: the adapter derives a stable UUIDv5 and keeps the original asproperties.envelopeId. Passing one straight toposthog-nodewould lose the dedupe silently. withDevFallback's[dev-fallback]marker is the same idea asdeduplicated: trueand theCANNOT FAILverdict — a run that did less than it looks like it did says so.- There is no
/healthto migrate. AddcreateHealthRegistrywith Postgres and Redis at tier 1 and the search provider at tier 3, and--provein the gate.
Consumable from CommonJS
A Medusa backend is module: Node16 CommonJS by upstream requirement, not by preference. The
library entries here ship both builds and a require condition carrying their own .d.cts, so a
static import type-checks and require() works:
import { createTracer } from '@geonosis/observability' // ESM
const { createTracer } = require('@geonosis/observability') // CommonJSgeonosis-observability, the bin, stays ESM alone: nothing require()s a bin, and its entry opens
with a top-level await that CommonJS cannot express. Building it twice would mean rewriting a
working entry point for a consumer that does not exist.
