npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@keyring-dev/sdk

v0.1.1

Published

The Keyring SDK runtime: the framework-agnostic request decision, the three onUnavailable modes, per-env resource injection, runInEnv, on-demand policy fill and the telemetry ring buffer. Framework adapters are thin wrappers over this.

Readme

@keyring-dev/sdk

The Keyring SDK runtime: the per-request decision, the three onUnavailable modes, per-environment resource injection, runInEnv, on-demand policy fill and the telemetry ring buffer. MIT, and it ships inside the customer's process.

Framework adapters — @keyring-dev/express, @keyring-dev/fastify — are thin wrappers over this. An adapter stays 40–80 lines, and that is only true if everything an adapter is tempted to reimplement lives here instead. It is also what makes "then Python, then Go" a week each: @keyring-dev/core is the pure function, this is the runtime, and the adapters are translation.

The hot path has no await in it

Keyring.handle() is synchronous from the bearer token to the decision. So is PolicyStore.get(), and so is status(). That is not a style preference: an async signature anywhere on this path makes a network call on it expressible, and the entire architecture rests on it not being. Verification is a SHA-256, a map lookup and a constant-time compare — measured at p50 3.2 µs with 10,000 keys cached.

req.keyring

The surface every integrator writes code against, and the one whose shape costs a major version across three languages to change.

The property is KeyringContext | null in every adapter — @keyring-dev/express and @keyring-dev/fastify spell it identically, and the Python and Go ports inherit that. It is null on a route configured skip: true and on a request the adapter refused, so a TypeScript caller narrows it; the examples below elide the narrowing to keep the field they are about in view.

The type is identical; the guarantee is not, and a port has to carry both halves. Fastify's holds unconditionally — decorateRequest puts the field on every request at creation, before any hook. Express's holds from the middleware onward — nothing decorates an Express request, so the adapter can only assign, and a request that never reached keyring() has no property at all. A framework whose port can decorate should; one that cannot must say where its guarantee starts. Read the property with ?., not with === null: under this type a !== null guard type-checks in a pre-mounted middleware and then throws.

| Field | Meaning | | ----------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | | keyId, workspaceId, projectId, tenantId | The identity. null on a fail-open decision, where we have no policy record. | | env | 'live' \| 'test', read from the key's own hashed material. | | scopes | What the key was minted with. Empty on a fail-open decision. | | displayPrefix | Loggable by construction; never a credential. | | expiresAt | Epoch ms, or null. | | degraded | True whenever the decision came from stale or incomplete policy — not only when a never-seen key was let through. | | verified | False when the key was admitted with no policy record at all. | | has(scope) | Scope matching, the same function the conformance fixtures pin. | | assertEnv(env) | Throws unless this request is that environment. | | your resources | Level 2, below. |

The three ergonomic levels

// Level 1 — read it.
app.get('/v1/orders', (req, res) => {
  const db = req.keyring.env === 'test' ? testDb : liveDb
})

// Level 2 — declare it once. This is the one the docs lead with.
const keyring = new Keyring({ resources: { db: { live: liveDb, test: testDb } } })
app.get('/v1/orders', (req, res) => req.keyring.db.orders.findMany(...))

// Level 3 — a scope for code that is far from the request.
await keyring.runInRequest(req.keyring, async () => {
  await sendReceiptEmail(order)   // reads currentEnv() from AsyncLocalStorage
})

Level 2 turns "remember to check env" from a discipline into a wiring decision made once. Level 3 is what makes background jobs and email senders env-aware without threading a parameter through six frames — this is the place where every hand-rolled test mode actually leaks.

A resource may not be named after a field of req.keyring; the constructor throws rather than letting resources: { env: ... } shadow the one field the whole test-mode story is read from.

onUnavailable

| Mode | A hit | A key the cache has never seen | | ----------------------------- | ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | | stale-then-open (default) | Served however stale, degraded set past the budget | Allowed, degraded: true, verified: false — unless the cache can prove the key does not exist, in which case 401 | | stale-then-closed | Served however stale | 401 | | closed | 503 past maxStalenessMs | 503 |

Why stale-then-open is defensible, and exactly how much it admits. The cache is complete for the project, not a partial memo, so a key absent from a fresh complete snapshot is a key that does not exist or was revoked, and the store says so: that miss is a 401 in every mode. The fail-open branch is reached only when the store cannot prove the miss, and what it admits then depends entirely on which of those states the store is in. It is not one number:

| Store state | What a miss admits | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Complete and fresh | Nothing. 401 — the store proved the key absent | | Complete, stale within maxStalenessMs | Still nothing — a miss is still proven, same as fresh. A revocation issued inside the window is the only thing not yet honoured on a hit (isAuthoritativeMiss reads only staleness and completeness, never a policy's own content) | | Complete, stale past maxStalenessMs | Any well-formed key — the store can no longer prove absence at all, so a miss falls through to the same admission as the four rows below it. Measured on a live rehearsal: a key generated with mintKey and never minted through the API answered 401 for 54.6 s, then 200 | | Never loaded (cold start, no snapshot) | Any well-formed key, for as long as the control plane stays unreachable or keeps answering 429/5xx — unbounded. A 400/403/404 here is the row below instead | | Evicted under maxCachedKeys (complete: false) | Any well-formed key not resident, for as long as the store stays incomplete | | Tenant-narrowed (tenantIds, complete: false) | Any well-formed key of any tenant outside the filter — permanently, by construction | | Snapshot walk cut short by the page bound (snapshotTruncated) | Any well-formed key the walk never reached. Only reachable from a control plane that answers a cursor with almost no keys; a project that is merely large no longer produces this state | | Auth-failed (the policy poll itself was refused with 401, or a cold node — one that has never loaded a snapshot — got a non-transient 400/403/404) | No unknown key, from the first refused poll until one succeeds: a miss is refused (401; 503 under closed, as before) whichever row the store would otherwise be in, "Never loaded" included. A hit still verifies from the last snapshot, so a key revoked after the first refused poll keeps verifying on this node. See below |

The first two rows are the case the default is designed for, and there a miss admits nothing at all — the store still proves it. The next five are the cases fail-open actually fires in, whether that is an ordinary outage past maxStalenessMs (60 s by default) or one of the four structural states below it, and there the radius is every well-formed string: the key format's checksum is CRC-32, a public algorithm, so producing one costs an attacker nothing. A tenant-narrowed node never leaves that state — no snapshot or delta on that filter will ever carry a tenant the filter excludes (see On-demand fill).

Every row above is built for a control plane that is slow or unreachable. A 401 on the policy poll says the opposite: this node's own credential no longer authenticates, most often because GitHub secret scanning revoked a leaked krsk_live_ key and it was this deployment's own KEYRING_SECRET_KEY, and a refused credential is not evidence that any other key is fine to admit. So from the first poll that comes back 401 until a poll succeeds, a miss is refused (401; 503 under closed, as before) regardless of maxStalenessMs and of onUnavailable, whichever row the store would otherwise be in. A poll that fails any other way (a timeout, a 5xx) leaves the state as it is. This is fail static, not refuse-all: a hit still verifies from the store exactly as it always does, which also means a revocation issued after the first refused poll does not reach this node until a poll succeeds. Rotating the secret key is what ends it.

A cold node's boot refusal on a 400/403/404 enters the same state. 401 is not the only status a boot-time misconfiguration produces: a krsk_test_ key naming env: 'live' gets 403, a wrong KEYRING_PROJECT_ID gets 404, and a malformed request gets 400. A 403 never clears on its own. A 404 or 400 usually does not either -- but can also be the status a refused tenant widening produces (a tenant_id this workspace does not own fails the whole snapshot the same way), and that cause is not a misconfiguration: the SDK's own rollback drops the offending id and the very next poll succeeds, with no operator action. The status code alone cannot tell the two apart, which is why PolicyBootRefusedError's message names both rather than asserting one. Before this, a cold node -- one that has never loaded a snapshot -- left in "Never loaded" on one of these statuses admitted any well-formed key for as long as the condition lasted, which for a genuine misconfiguration is the life of the deploy. So the same override now applies, scoped to the cold refusal only: a warm node's ordinary failed poll on one of these statuses is unchanged, because it has a known key to fall back on and a control plane that starts refusing everything after having served correctly is a different failure from a boot that never authenticated at all. 429 and every 5xx are excluded on purpose -- they are transient, and a cold node keeps admitting any well-formed key while the control plane is merely slow or unreachable, same as before. onError receives a PolicyBootRefusedError rather than a PolicyAuthFailedError for this cause, because the likely reason and the fix differ; stats().policyCache.authFailed and authFailedRefusingAll read the same either way.

Two cases sit outside the state. A node that starts or restarts is in its own row above until its first poll comes back: with no disk snapshot, or with one older than maxStalenessMs, it admits any well-formed key for that first round trip. And a node built with store has no poller, so it never enters the state.

stats().policyCache.authFailed reports the state, distinct from ordinary staleness. onError receives a PolicyAuthFailedError or a PolicyBootRefusedError at most once a minute: on the refused poll that enters the state, unless one already went out in the minute before (so a credential flapping between 401 and success does not re-announce every time), and again while the state continues. That is on top of the raw PolicyRequestError every failed poll already reports.

A node that has never loaded a snapshot at all and gets a 401, or one of these 400/403/404 statuses, refuses every key, because it has no known key to fall back on; drain it the way you drain scopeTooLarge:

app.get('/readyz', (_req, res) => {
  const { authFailedRefusingAll, scopeTooLarge } =
    keyring.stats().policyCache ?? {};
  res
    .status(
      authFailedRefusingAll === true || scopeTooLarge === true ? 503 : 200,
    )
    .end();
});

Out of scope, deliberately, because this change does not touch either path. The rate limiter (/v1/ratelimit/check) and idempotency (/v1/idempotency/*) authenticate with the same credential, so a revoked key 401s them too — but each keeps its own existing rateLimitOnUnavailable (open by default) and idempotency onUnavailable (closed by default) behaviour for that failure, unchanged by the auth-failed state above.

A tenant the node has narrowed to can also be temporarily outside the filter, while a refused widening is adjudicated, and that is the same row for as long as it lasts. How long it lasts is measured, not promised: on a node with no traffic of its own, 50 s for a burst carrying one bad id, 420 s for four, 790 s for eight, 1 150 s for sixteen and 1 725 s for thirty-two; 50 s and then 85–115 s on a node whose request path re-offers what it cannot verify. A tenant the node has convicted — two refusals, at least one naming it alone — is a permanent case rather than a temporary one on a quiet node. See On-demand fill for the mechanism and the shorter, filter-membership figure it is stated in.

A project larger than maxCachedKeys used to be a fifth such row. Every re-seed stopped at the same bound, complete stayed false for the life of the process, and the node admitted any well-formed key for as long as it ran. Up to maxAutoCachedKeys that row is gone: the bound raises itself, on every path a project crosses it on. Past maxAutoCachedKeys it is still that row on a node that is already serving — the node keeps serving because killing it is a fleet-wide outage of your own API, and it says so loudly on every poll rather than silently. See A project this node cannot cache.

What the branch does not do is invent an identity. Every id is null, scopes is [], has() returns false, verified is false and degraded is true; a node pinned with env still refuses the other environment, a revoked key stays revoked on a stale store, and a krsk_ secret key presented as a tenant key is refused. degraded and verified are the whole of what separates this branch from an authentication bypass, so:

  • alert on degraded, and on stats().policy.fetchedAt not advancing;

  • use closed (or stale-then-closed) on any route where admitting an unknown key is unacceptable — money routes are the obvious ones;

  • in a containerised deployment, reach for stale-then-closed first. A fresh pod has no snapshot to be stale from, so it sits in the third row — every well-formed key admitted unverified — for the whole outage, and no volume configuration removes that for a pod that is born without a file. The cost is stated rather than softened: during an outage such a pod refuses every request with a 401 until its first successful poll. That is an availability cost against an authentication one, and it is the customer's call; stale-then-open remains the default.

  • ship a disk snapshot (persist, on by default) so a restart that keeps its filesystem is the second row rather than the third; poller.start() hydrates from it before its first poll for exactly this reason. This does not cover an immutable-container rollout, and mounting a volume is the second remedy there rather than the first: it narrows the window instead of closing it, emptyDir buys nothing, a default rolling update and a scale-up are not covered by a ReadWriteOnce PVC or a volumeClaimTemplate, and a pod that does hydrate a pre-outage snapshot is the "Complete, stale" row — it serves, but revocations issued during the outage are not honoured. See @keyring-dev/cache's README, "Immutable containers and rollouts".

    The second row is the best case, not the only one. It assumes the hydrated file is one the node may believe. A file written before the SDK recorded its scope, or by a tenant-narrowed process, loads incomplete on purpose — the "Evicted … complete: false" row — so until that node's first repair walk lands it verifies every key the file holds and admits any well-formed key that is not resident, including a forged one, where a node that (wrongly) believed the file would have answered 401. Bounded by maxStalenessMs: past it a complete-but-stale store cannot prove a miss either.

  • a node refuses a disk snapshot written under a different scope — a tenant-narrowed process must not hydrate a project-scoped process's file — and starts cold instead, reporting it once through onError. Give each scope its own cacheDir when several processes share one filesystem.

verified: false also refuses idempotency outright, because there is no scope to key a record on — see A request we cannot scope.

That argument is load-bearing, so the code honours it rather than assuming it. PolicyStore.status().complete is false whenever the store cannot prove a miss — after an eviction under maxCachedKeys, when the snapshot was narrowed to some tenants, or when the poller stopped a paged snapshot walk before its last page — and every miss from such a store lands in the degraded branch instead of the 401 one.

A project this node cannot cache

A node whose project holds more live keys than its maxCachedKeys cannot ever build a complete cache, and an incomplete cache under the default above admits every well-formed key, unverified, for the life of the process. Documenting that is not the same as choosing it, so the SDK does two things in order.

First it raises maxCachedKeys by itself, up to maxAutoCachedKeys (38,000 by default). Measured on the path the poller actually takes — a page of snapshot JSON through JSON.parse and into the store — the cost depends on the shape of your keys rather than on their number:

| shape | bytes/key | at the 38,000 ceiling | | ------------------------------ | --------- | --------------------- | | 1 scope, no rate limits | 519 | 18.8 MiB | | 3 scopes, a two-rule limit set | 855 | 31.0 MiB | | 8 scopes, a five-rule set | 1 271 | 46.0 MiB |

against ~8 MiB at the 10,000-key default on the middle row. A re-seed costs about 1.4× more than the steady state while the page in flight is still uncollected, and one PATCH /v1/projects/:id on your project's default rate limits makes every node in your fleet re-seed at the same moment — so size the container for the peak, not the steady state. The raise is reported through onError as a PolicyCacheAutoRaised — a notice, not a failure; the poll it came from succeeded — and it is scrapeable:

keyring.stats().policyCache;
// { autoRaisedTo: 38000, maxAutoCachedKeys: 38000,
//   snapshotPages: 4, snapshotTruncated: false, scopeTooLarge: false,
//   diskSnapshotRefused: false, authFailed: false, authFailedRefusingAll: false }
// autoRaisedTo is null on a node that never had to grow.
// diskSnapshotRefused is true on a node that found a disk snapshot for its
// project and environment and refused it for being another scope's -- the
// difference between "no file" and "a file this node must not read".
// stats().cache.configuredMaxCachedKeys is what you asked for; .maxCachedKeys
// is what is in force now.

If the project still does not fit, you get a PolicyScopeTooLargeError, and what it does depends on whether this node has ever served:

| the node | what happens | | ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | has never served (cold) | Refuses to start. The store is never loaded, the poller stops, and the SDK rethrows out of a microtask — an uncaught exception, so the deploy dies | | is already serving (warm) | Keeps serving and keeps polling. Records the condition, reports it through onError on every poll, and sets stats().policyCache.scopeTooLarge |

The asymmetry is deliberate and it is the whole point. "Refuse to start" and "kill a process that is answering requests" are different actions: your API runs in the same process, and a floor raise re-seeds a whole fleet in the same second, so throwing at a warm node is a fleet-wide outage of your product from a routine action on ours. Stopping its poller would be worse still — a node that stops polling never learns another revocation.

Drain a warm node rather than letting the SDK kill it. scopeTooLarge is a readiness predicate:

app.get('/readyz', (_req, res) => {
  const { scopeTooLarge } = keyring.stats().policyCache ?? {};
  res.status(scopeTooLarge === true ? 503 : 200).end();
});

Your orchestrator then stops routing to it and fails the deploy, which is what "refuse to start" means to an SRE, without killing a process that is still answering. The condition clears itself when a later walk finds the project fits again, so raising the bound recovers the node on its next re-seed.

The fix, in the order the message names it. Raise maxAutoCachedKeys (or maxCachedKeys, which the ceiling never undercuts) and pay ~855 bytes of heap per cached key. Setting maxAutoCachedKeys equal to maxCachedKeys turns the automatic step off and goes straight to the refusal, which is the right shape for a node with a hard memory bound.

new Keyring({ maxAutoCachedKeys: 100_000 }); // ~82 MiB on the middle row above

tenantIds is not that fix, and the error says so in the same sentence that offers it. Narrowing lets the node start, and a tenant-narrowed store is complete: false by construction — see the table above — so every key outside the filter is admitted unverified under the stale-then-open default, permanently. That is the state the refusal exists to prevent, reached by following a remedy. Narrow only with stale-then-closed or closed on the routes where admitting an unknown key is unacceptable.

onScopeTooLarge takes the cold decision back — to page, to fail a readiness probe, to exit with your own code. It is not called on a warm node. A no-op there leaves the node serving on the store it has, which at boot is an empty one: every well-formed key admitted unverified, with no successful poll coming.

When it fires. On a walk — a cold start, a restart with no readable disk snapshot, or a delta the feed answers with a snapshot — and on every other way a project-scoped store loses completeness, because that is the state, not the walk. A delta that admits a key past the bound (ordinary growth: a project crosses maxCachedKeys by minting keys, with nothing re-seeding), a disk snapshot an earlier process wrote short and start() hydrates from before its first poll, and a last page larger than the max_keys this client asked for all re-seed the same way. persist is on by default, so for a restart that keeps its filesystem the disk path is the ordinary restart, and it is a delta rather than a walk — a fresh pod in an immutable-container rollout has no file to hydrate and re-seeds with a full walk instead, same as any other cold node.

Two cases are deliberately not touched. A tenant-narrowed node that still does not fit stays a truncation: its store is complete: false by construction already, so refusing to serve would buy nothing the tenant filter had not already cost. And a walk cut short by the page bound (MAX_SNAPSHOT_PAGES) stays a truncation too: that bound is about a control plane answering a cursor with almost no keys, and a bug on our side must not be able to stop a customer's fleet.

The sharp edge: a route that declares scopes is not scope-checked in the fail-open branch, because checking against a fabricated empty scope set would be inventing an answer rather than admitting we have none. Use stale-then-closed (or closed) on a route where that is unacceptable:

new Keyring({
  routes: {
    'POST /v1/payments': { onUnavailable: 'closed' }, // money: refuse rather than guess
    'GET  /v1/health': { skip: true },
    '/v1/internal/*': { onUnavailable: 'stale-then-closed' },
  },
});

Route keys are METHOD /path, /path or /prefix/*; the most specific rule wins regardless of declaration order, and keyring.stats().unmatchedRoutes names any rule that has never matched — a typo in a route key is otherwise a silent downgrade to the fail-open default.

Rate limiting

Limits are policy, and policy already has a distribution channel. A key's effective limits — its own override, or the project default, resolved by the control plane so the layering is in one place and not in three SDKs — arrive on the same snapshot and delta feed the keys do.

That is what keeps the round trip off most requests: a key with no limits never talks to us at all. handle() stays synchronous; the check is a separate awaited step the adapter makes only when a limit actually applies.

The counters are central, in our Redis, reached over the same krsk_ credential and the same base URL the policy poller already uses — never by handing your process Redis credentials. Every limit in a check is evaluated in one round trip, and all the counters are read before any of them is charged, so a caller over their daily quota does not keep burning their per-second one.

Two algorithms, and the choice is about semantics rather than speed (all four candidates measured within 11 % of each other):

| Limit shape | algorithm | Why | | --------------------------------- | ----------- | ------------------------------------------------------------------------------------------ | | N per second or minute | sliding | Current window plus a weighted slice of the previous one. No boundary burst. | | N per day or per billing period | fixed | What a customer means by a quota. A boundary burst at day granularity is nobody's problem. |

scope: 'tenant' shares a counter across every key that tenant holds, which is what "this organisation gets 10,000 a day" means and what survives a rotation. scope: 'key' is per key.

Headers

Both spellings ship, and the legacy triple is on by default because every client library in production parses it.

RateLimit-Policy: "per-key-second";q=10;w=1, "per-tenant-day";q=10000;w=86400
RateLimit: "per-key-second";r=7;t=1, "per-tenant-day";r=9000;t=3600
RateLimit-Limit: 10          # the most constrained policy, not the first
RateLimit-Remaining: 7
RateLimit-Reset: 1
Retry-After: 1                            # on a 429, never before the window rolls
Keyring-Rate-Limited-Policy: per-key-second

Set legacyRateLimitHeaders: false to emit only the current draft.

When the counters are unreachable

rateLimitOnUnavailable — open by default.

| Mode | What happens | | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | | open | Served, degraded set, no RateLimit header at all — we do not know the numbers, and a fabricated remaining is a number your caller will act on. | | closed | 503 with Retry-After. For a route where admitting unmetered traffic is the expensive failure: an LLM call, an outbound SMS, anything you pay per unit for. |

There is no stale-then-open here, because there is no stale state to serve from: a counter is a shared number, not a cached record. A local one is a different limiter — a static split, measured at +12 % overshoot and −6 % under-delivery — and v1 does not ship it.

The wait is bounded (rateLimitTimeoutMs, 500 ms). Our availability is not yours, and an unbounded wait on our socket is the one line that would make that false.

Idempotency

Stripe's semantics, because every backend engineer already knows them: header Idempotency-Key, up to 255 characters; the status and body of the first request are stored whether it succeeded or not, so a replay of a request that 500'd returns the same 500; a reused key carrying a different request is a 422; a duplicate arriving while the first is still running is a 409 with Retry-After rather than a block.

Two deliberate deviations, and one more below:

  1. PATCH and PUT are covered too. "PUT is idempotent by definition" is true of the resource state and false of the email it sends.
  2. TTL is configurable, 24 h by default and 7 days at most, because some clients retry from a dead-letter queue the next business day.

Scope is per (project, env, tenant), not per key. Per-key scoping breaks rotation-with-overlap: the old key writes the record, the retry arrives on the new key, and your customer is charged twice by the feature that exists to prevent exactly that.

Mounting it

The fingerprint covers the method, path, sorted query and body — never the headers, because a retry legitimately carries a different User-Agent or trace header. It therefore needs a parsed body, which is why it is a second Express mount, after your body parser:

const mw = keyring({ secretKey, projectId });
app.use(mw); // verify + rate limit, no body needed
app.use(express.json());
app.use(keyringIdempotency(mw.keyring)); // idempotency, after the parser

Fastify needs no second registration: the plugin hooks onRequest for verification and the limit, and preHandler for idempotency.

The body is canonicalised before hashing — object keys sorted, arrays left alone — so a retry serialised by a different client library is the same request.

At-least-once, and we cannot fix it

If your process dies after your handler has committed but before the record is written, the retry re-executes. The record is left in_flight, the retry finds an expired lock, steals it and runs your handler again. This is genuine at-least-once behaviour, it is a property of every middleware-shaped idempotency layer, and the only real fix is a transaction spanning our store and your database, which we do not have.

The mitigation is the standard one: make your handler's own write idempotent — a unique constraint on your side, keyed on something derived from the request — and pass the same key through. A process that dies while the response is streaming is fine: the record is written before the body is flushed.

A legitimately slow handler is fine too. The 30 s lease is refreshed by a heartbeat while your handler runs, so a 45 s handler is not double-executed by an eager retry.

Size, and when the store is unreachable

Responses over 256 KB are not stored. The record is kept, so conflict detection still works, and a replay answers 409 response_too_large_to_replay telling the client to check the resource — never a wrong body.

idempotency.mode — closed by default, which is the opposite of the rate limiter's, and the asymmetry is the decision. A rate limit not enforced during an outage admits excess traffic; an idempotency record not written executes a payment twice, and the request in front of us has explicitly asked for exactly-once by carrying the header. Refusing it with 503 is declining to promise what we cannot deliver, to a caller that is holding a retry loop and will ask again. A request with no Idempotency-Key never reaches this code, so the blast radius is only the mutating requests whose clients were built to retry.

mode: 'open' serves them anyway, for a customer whose handlers are already idempotent on their own side. idempotency: false removes the feature.

A 5xx your transient predicate accepts — by default Retry-After present, or status 503 — releases the record instead of storing it, so the retry executes. Their 503 is usually "my database was failing over", not "this operation was attempted".

Only a 5xx. The predicate is consulted for status >= 500 and for nothing else, whatever it returns, because a 2xx or a 4xx is an attempt that happened and releasing its record re-executes the operation. 202 Accepted + Retry-After + Location is the standard async-job convention, and a job started by a POST is exactly the mutating retried request this feature exists for — without the gate the retry started a second one.

A request we cannot scope

A request whose decision was not verified is refused idempotency. req.keyring.verified is false on the stale-then-open branch — before this node's first snapshot lands, after an eviction, past maxStalenessMs — and on that branch we know the environment (it is inside the key's own hashed material) and neither the project nor the tenant. There is no (project, env, tenant) scope to key a record on, so closed answers 503 and open serves the request with no exactly-once promise, exactly as each mode does everywhere else.

It is the same judgement closed already encodes: we do not promise exactly-once when we cannot deliver it. Scoping the record on something key-specific instead would be worse — the retry that arrives once the snapshot has landed is verified, addresses the real scope, finds nothing, and re-executes having been told it would not. keyring.stats().idempotency.unverified counts these separately from unavailable: our store was reachable, the policy was not.

On-demand fill

A miss the store cannot prove schedules an asynchronous policy refresh. The request that triggered it never waits: a "miss → single remote verify" design was deliberately rejected, because that branch is exactly what puts a round trip on a cold cache. The fill only makes the next request better.

It is bounded three ways — one fetch in flight, a floor between fetches, a ceiling per minute — because a scan of invented keys must not become a scan of our own control plane.

A node configured with tenantIds is the one case a refresh cannot repair: no delta or snapshot on that filter will ever carry a tenant the filter excludes, and deriving the tenant from an unknown key hash would mean asking us on the hot path. Such a node calls keyring.admitTenants([...]) — the caller knows which of their tenants is calling; the SDK does not.

The ids are Keyring tenant ids (the UUIDs the control plane minted), not your own external_id. An id that is not one is ignored, and every id a refused snapshot carried is rolled back out of the filter — the node keeps serving from the scope it already had. Without both, a single typo in a vendor's externalId → keyringTenantId mapping left the node's snapshot 404ing forever: stale past maxStalenessMs, admitting every miss under the fail-open default, and receiving no revocation. stats().policyPollFailures is the number to alert on.

A refusal names the filter, though, not an id — a 404 (or a 400: the node treats them identically) is equally what a rolling deploy, a proxy or a gateway answers — so the rollback is a suspicion and not a verdict. One refusal holds every id that request carried, out of the filter and against admitTenants, for a cooldown. The node re-seeds on the scope that worked and then re-admits the suspects by bisection: half of a refused set goes back in, a poll that serves it acquits that half outright, and the search narrows to the other.

What two refusals buy is the opposite of a hold: an id refused twice, at least once on a request that named it alone, is convicted and leaves the cooldown's protection for good — the node stops offering it, and only the customer's own admitTenants or a restart brings it back. One refusal is not enough for that, even alone, because the rolling deploy above answers the same 404: a recurring one that landed on a probe of a single id used to convict the legitimate tenant it named. Two of them still can. If your tenant list is machine-generated and your control plane 404s intermittently, a legitimate tenant can be convicted and stay out until the process restarts; on a node with traffic for that tenant the next cache miss restores it, on a quiet one nothing does.

This matters because a tenant-narrowed store is not complete, so while a tenant is out of the filter every request it makes is admitted unverified under the fail-open default (or refused outright under closed): time out of the filter is the cost, and re-admitting one id per poll made it grow with the size of the refused burst — 50 tenants admitted alongside one bad id waited 50 polls.

There are two windows here and they are not the same length. The one that matters to you is how long the tenant's keys are not in the store, because that is how long its traffic is admitted unverified under the fail-open default — or refused outright under closed. Measured on a 50-tenant burst, on a node with no traffic of its own: 50 s with one bad id in it, 420 s with four, 790 s with eight, 1 150 s with sixteen and 1 725 s with thirty-two — it grows with the number of ids your mapping got wrong. On a node whose request path re-offers the tenants it cannot verify it is 50 s, and then 85–115 s at every size. That is the number this mechanism should be judged on and the one to size onUnavailable against.

The shorter window is membership of the filter, and it is the mechanism's own figure rather than yours: no id the node takes out of the filter stays out for more than one cooldown plus the poll that notices it is due, which is 60 s plus up to 30 s at the defaults: 322 of 366 swept shapes land in the 61.5–69.9 s band, on this revision and on every one before it. When the cooldown runs out the node puts the id back regardless of what the search has concluded, which is exactly what it did before there was a search at all. The "plus the poll" is not a hedge: a poll following a failed one backs off to as much as six intervals. Being in the filter for the duration of one request is not the same as being verifiable, which is why the two numbers differ by up to twenty times on the bursts above — and why the customer-facing one leads.

Two cases fall outside even that bound and both are earned: past MAX_HELD_TENANTS there is no record left to put back with, and the id is merely admittable again — the same thing, one step weaker — and an id the search convicts leaves the bound for good, so that the node stops offering a filter the control plane demonstrably refuses once a minute for ever. A conviction takes two refusals for exactly that reason.

Putting an id back in the filter and putting it back on the wire are two different acts, and only the first is on that clock. A poll that follows a failed poll never carries a widening the control plane has not served: it fetches a delta instead, which carries no tenant_id at all and so cannot be refused for the filter. So the store stays fresh and the search keeps running while the widening waits for the one poll it takes to get a success behind it. Without that split the two properties really are incompatible — a snapshot carries the whole filter, so re-admitting on the clock alone meant a dozen refused ids arriving on a dozen polls could put a refusal on every poll for ever, and only a restart cleared it. On top of that the search is best effort: one bad id inside a thousand is convicted in 10–25 polls rather than 1,024, and a burst carrying several is measured rather than promised, because separating K good ids from d bad ones against a yes/no answer costs at least log₂(C(K+d, d)) experiments and no implementation beats that. stats().tenantsHeldOut is how many are out, counting the one the node is currently probing; it is bounded, so a caller passing a fresh UUID per request cannot grow it without limit — and once that bound binds, widening pauses for a cooldown, all widening, including tenants that have never been refused, because past the cap an evicted id and a new one are the same thing to this node.

admitTenants returns false for three different reasons — the id is already in the filter, this node was never narrowed, or the value is not a tenant id — so the one you have to act on has its own signal: a value dropped for shape is counted in stats().tenantIdsRejected and reported to onError. No poll fails in that case, so nothing else moves.

Telemetry

Fixed 10,000-event ring buffer (~2.5 MB), drop oldest when full, count the drops, warn once a minute. Flush every 1,000 ms or 1,000 events, gzipped, fire-and-forget with a 2 s timeout; one retry with jitter, then drop — we are not a log shipper. A 429 is never retried: the Retry-After is honoured, because retrying into a rate limit is the storm the limit exists to survive. Above maxEventsPerSecond (5,000) non-error events are sampled and the rate is recorded in the batch, but every 4xx, 5xx and degraded event is kept — the customer debugs errors, not successes.

On SIGTERM and beforeExit the buffer is flushed with a 3 s cap. Adding a SIGTERM listener suppresses Node's default action, so the handler re-raises the signal once the flush is done and only when nothing else was listening — a telemetry flush that quietly made a customer's containers un-killable would be a far worse bug than losing a batch.

The query string is stripped from every path before an event leaves the process: a vendor whose callers pass a token in a query parameter would otherwise have it shipped to us and stored in our request log.

Log redaction

The single most common way an API key product leaks keys is its own customers' logs. redact(value) deep-clones any value and masks a Keyring key (kr_/krsk_) or an embed token wherever it sits — a bare string, a header, a nested object, an Error's message — plus any field whose name looks sensitive (password, secret, authorization, ...), whatever that field's shape:

import { redact } from '@keyring-dev/sdk';

logger.info(redact({ headers: req.headers, err }));

Pre-built redactors wrap this for pino, winston and bunyan, because none of the three offer a pattern-based hook on their own:

import {
  pinoRedact,
  winstonRedact,
  bunyanRedactStream,
} from '@keyring-dev/sdk';

const logger = pino({ hooks: pinoRedact() });

pinoRedact() returns both logMethod (masks the arguments passed to a log call) and streamWrite (masks the final serialized line, which is the only hook pino runs after its own req/res serializers and after a child logger's bound fields are merged in — pass both, as hooks: pinoRedact() does, or pino-http's request line and a child logger's bindings go out unmasked).

format.combine(format(winstonRedact())(), format.json());
bunyan.createLogger({
  name: 'app',
  stream: bunyanRedactStream(process.stdout),
});

bunyanRedactStream sees every line bunyan writes, whatever field it is under, and is the one that gives full coverage. bunyanRedactSerializers covers only the field names it names (err, error, req, res, key) — cheaper when other bunyan streams already exist and should not be double-wrapped, but a key under a field name it does not list goes out unmasked; use the stream wrapper unless that trade is one you want to make.

redactString(value) is the same scan over one string, for a log line built by hand rather than through a logger's own call.

Configuration

rateLimitOnUnavailable (open), legacyRateLimitHeaders (true), rateLimitTimeoutMs (500), idempotency (true, or an options object, or false).

secretKey, projectId and baseUrl default to KEYRING_SECRET_KEY, KEYRING_PROJECT_ID and KEYRING_BASE_URL. A kr_ tenant key passed as secretKey is refused: it cannot authenticate, and sending it would be the customer's own credential leaving their process by mistake.