@keyring-dev/sdk
v0.1.1
Published
The Keyring SDK runtime: the framework-agnostic request decision, the three onUnavailable modes, per-env resource injection, runInEnv, on-demand policy fill and the telemetry ring buffer. Framework adapters are thin wrappers over this.
Readme
@keyring-dev/sdk
The Keyring SDK runtime: the per-request decision, the three onUnavailable
modes, per-environment resource injection, runInEnv, on-demand policy fill and
the telemetry ring buffer. MIT, and it ships inside the customer's process.
Framework adapters — @keyring-dev/express, @keyring-dev/fastify — are thin wrappers
over this. An adapter stays 40–80 lines, and that is only true if
everything an adapter is tempted to reimplement lives here instead.
It is also what makes "then Python, then Go" a week each: @keyring-dev/core is the
pure function, this is the runtime, and the adapters are translation.
The hot path has no await in it
Keyring.handle() is synchronous from the bearer token to the decision. So is
PolicyStore.get(), and so is status(). That is not a style preference: an
async signature anywhere on this path makes a network call on it expressible,
and the entire architecture rests on it not being. Verification is a SHA-256, a
map lookup and a constant-time compare — measured at p50 3.2 µs with 10,000 keys
cached.
req.keyring
The surface every integrator writes code against, and the one whose shape costs a major version across three languages to change.
The property is KeyringContext | null in every adapter — @keyring-dev/express
and @keyring-dev/fastify spell it identically, and the Python and Go ports inherit
that. It is null on a route configured skip: true and on a request the
adapter refused, so a TypeScript caller narrows it; the examples below elide the
narrowing to keep the field they are about in view.
The type is identical; the guarantee is not, and a port has to carry both
halves. Fastify's holds unconditionally — decorateRequest puts the field on
every request at creation, before any hook. Express's holds from the
middleware onward — nothing decorates an Express request, so the adapter can
only assign, and a request that never reached keyring() has no property at
all. A framework whose port can decorate should; one that cannot must say where
its guarantee starts. Read the property with ?., not with === null: under
this type a !== null guard type-checks in a pre-mounted middleware and then
throws.
| Field | Meaning |
| ----------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- |
| keyId, workspaceId, projectId, tenantId | The identity. null on a fail-open decision, where we have no policy record. |
| env | 'live' \| 'test', read from the key's own hashed material. |
| scopes | What the key was minted with. Empty on a fail-open decision. |
| displayPrefix | Loggable by construction; never a credential. |
| expiresAt | Epoch ms, or null. |
| degraded | True whenever the decision came from stale or incomplete policy — not only when a never-seen key was let through. |
| verified | False when the key was admitted with no policy record at all. |
| has(scope) | Scope matching, the same function the conformance fixtures pin. |
| assertEnv(env) | Throws unless this request is that environment. |
| your resources | Level 2, below. |
The three ergonomic levels
// Level 1 — read it.
app.get('/v1/orders', (req, res) => {
const db = req.keyring.env === 'test' ? testDb : liveDb
})
// Level 2 — declare it once. This is the one the docs lead with.
const keyring = new Keyring({ resources: { db: { live: liveDb, test: testDb } } })
app.get('/v1/orders', (req, res) => req.keyring.db.orders.findMany(...))
// Level 3 — a scope for code that is far from the request.
await keyring.runInRequest(req.keyring, async () => {
await sendReceiptEmail(order) // reads currentEnv() from AsyncLocalStorage
})Level 2 turns "remember to check env" from a discipline into a wiring decision
made once. Level 3 is what makes background jobs and email senders env-aware
without threading a parameter through six frames — this is the place where
every hand-rolled test mode actually leaks.
A resource may not be named after a field of req.keyring; the constructor
throws rather than letting resources: { env: ... } shadow the one field the
whole test-mode story is read from.
onUnavailable
| Mode | A hit | A key the cache has never seen |
| ----------------------------- | ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| stale-then-open (default) | Served however stale, degraded set past the budget | Allowed, degraded: true, verified: false — unless the cache can prove the key does not exist, in which case 401 |
| stale-then-closed | Served however stale | 401 |
| closed | 503 past maxStalenessMs | 503 |
Why stale-then-open is defensible, and exactly how much it admits. The
cache is complete for the project, not a partial memo, so a key absent from a
fresh complete snapshot is a key that does not exist or was revoked, and the
store says so: that miss is a 401 in every mode. The fail-open branch is reached
only when the store cannot prove the miss, and what it admits then depends
entirely on which of those states the store is in. It is not one number:
| Store state | What a miss admits |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Complete and fresh | Nothing. 401 — the store proved the key absent |
| Complete, stale within maxStalenessMs | Still nothing — a miss is still proven, same as fresh. A revocation issued inside the window is the only thing not yet honoured on a hit (isAuthoritativeMiss reads only staleness and completeness, never a policy's own content) |
| Complete, stale past maxStalenessMs | Any well-formed key — the store can no longer prove absence at all, so a miss falls through to the same admission as the four rows below it. Measured on a live rehearsal: a key generated with mintKey and never minted through the API answered 401 for 54.6 s, then 200 |
| Never loaded (cold start, no snapshot) | Any well-formed key, for as long as the control plane stays unreachable or keeps answering 429/5xx — unbounded. A 400/403/404 here is the row below instead |
| Evicted under maxCachedKeys (complete: false) | Any well-formed key not resident, for as long as the store stays incomplete |
| Tenant-narrowed (tenantIds, complete: false) | Any well-formed key of any tenant outside the filter — permanently, by construction |
| Snapshot walk cut short by the page bound (snapshotTruncated) | Any well-formed key the walk never reached. Only reachable from a control plane that answers a cursor with almost no keys; a project that is merely large no longer produces this state |
| Auth-failed (the policy poll itself was refused with 401, or a cold node — one that has never loaded a snapshot — got a non-transient 400/403/404) | No unknown key, from the first refused poll until one succeeds: a miss is refused (401; 503 under closed, as before) whichever row the store would otherwise be in, "Never loaded" included. A hit still verifies from the last snapshot, so a key revoked after the first refused poll keeps verifying on this node. See below |
The first two rows are the case the default is designed for, and there a miss
admits nothing at all — the store still proves it. The next five are the cases
fail-open actually fires in, whether that is an ordinary outage past
maxStalenessMs (60 s by default) or one of the four structural states below
it, and there the radius is every well-formed string: the key format's
checksum is CRC-32, a public algorithm, so producing one costs an attacker
nothing. A tenant-narrowed node never leaves that state — no snapshot or delta
on that filter will ever carry a tenant the filter excludes (see
On-demand fill).
Every row above is built for a control plane that is slow or unreachable. A
401 on the policy poll says the opposite: this node's own credential no
longer authenticates, most often because GitHub secret scanning revoked a
leaked krsk_live_ key and it was this deployment's own
KEYRING_SECRET_KEY, and a refused credential is not evidence that any other
key is fine to admit. So from the first poll that comes back 401 until a
poll succeeds, a miss is refused (401; 503 under closed, as before)
regardless of maxStalenessMs and of onUnavailable, whichever row the store
would otherwise be in. A poll that fails any other way (a timeout, a 5xx)
leaves the state as it is. This is fail static, not refuse-all: a hit still
verifies from the store exactly as it always does, which also means a
revocation issued after the first refused poll does not reach this node until
a poll succeeds. Rotating the secret key is what ends it.
A cold node's boot refusal on a 400/403/404 enters the same state.
401 is not the only status a boot-time misconfiguration produces: a
krsk_test_ key naming env: 'live' gets 403, a wrong
KEYRING_PROJECT_ID gets 404, and a malformed request gets 400. A 403
never clears on its own. A 404 or 400 usually does not either -- but can
also be the status a refused tenant widening produces (a tenant_id this
workspace does not own fails the whole snapshot the same way), and that
cause is not a misconfiguration: the SDK's own rollback drops the offending
id and the very next poll succeeds, with no operator action. The status code
alone cannot tell the two apart, which is why PolicyBootRefusedError's
message names both rather than asserting one. Before this, a cold
node -- one that has never loaded a snapshot -- left in "Never loaded" on one
of these statuses admitted any well-formed key for as long as the condition
lasted, which for a genuine misconfiguration is the life of the deploy. So
the same override now applies, scoped to the cold refusal only: a warm
node's ordinary failed poll on one of these statuses is unchanged, because it
has a known key to fall back on and a control plane that starts refusing
everything after having served correctly is a different failure from a boot
that never authenticated at all. 429 and every 5xx are excluded on
purpose -- they are transient, and a cold node keeps admitting any
well-formed key while the control plane is merely slow or unreachable, same
as before. onError receives a PolicyBootRefusedError rather than a
PolicyAuthFailedError for this cause, because the likely reason and the fix
differ; stats().policyCache.authFailed and authFailedRefusingAll read the
same either way.
Two cases sit outside the state. A node that starts or restarts is in its own
row above until its first poll comes back: with no disk snapshot, or with one
older than maxStalenessMs, it admits any well-formed key for that first
round trip. And a node built with store has no poller, so it never enters
the state.
stats().policyCache.authFailed reports the state, distinct from ordinary
staleness. onError receives a PolicyAuthFailedError or a
PolicyBootRefusedError at most once a minute: on the refused poll that
enters the state, unless one already went out in the minute before (so a
credential flapping between 401 and success does not re-announce every
time), and again while the state continues. That is on top of the raw
PolicyRequestError every failed poll already reports.
A node that has never loaded a snapshot at all and gets a 401, or one of
these 400/403/404 statuses, refuses every key, because it has no known
key to fall back on; drain it the way you drain scopeTooLarge:
app.get('/readyz', (_req, res) => {
const { authFailedRefusingAll, scopeTooLarge } =
keyring.stats().policyCache ?? {};
res
.status(
authFailedRefusingAll === true || scopeTooLarge === true ? 503 : 200,
)
.end();
});Out of scope, deliberately, because this change does not touch either
path. The rate limiter (/v1/ratelimit/check) and idempotency
(/v1/idempotency/*) authenticate with the same credential, so a revoked
key 401s them too — but each keeps its own existing rateLimitOnUnavailable
(open by default) and idempotency onUnavailable (closed by default)
behaviour for that failure, unchanged by the auth-failed state above.
A tenant the node has narrowed to can also be temporarily outside the filter, while a refused widening is adjudicated, and that is the same row for as long as it lasts. How long it lasts is measured, not promised: on a node with no traffic of its own, 50 s for a burst carrying one bad id, 420 s for four, 790 s for eight, 1 150 s for sixteen and 1 725 s for thirty-two; 50 s and then 85–115 s on a node whose request path re-offers what it cannot verify. A tenant the node has convicted — two refusals, at least one naming it alone — is a permanent case rather than a temporary one on a quiet node. See On-demand fill for the mechanism and the shorter, filter-membership figure it is stated in.
A project larger than maxCachedKeys used to be a fifth such row. Every
re-seed stopped at the same bound, complete stayed false for the life of the
process, and the node admitted any well-formed key for as long as it ran. Up to
maxAutoCachedKeys that row is gone: the bound raises itself, on every path a
project crosses it on. Past maxAutoCachedKeys it is still that row on a node
that is already serving — the node keeps serving because killing it is a
fleet-wide outage of your own API, and it says so loudly on every poll rather
than silently. See
A project this node cannot cache.
What the branch does not do is invent an identity. Every id is null,
scopes is [], has() returns false, verified is false and degraded is
true; a node pinned with env still refuses the other environment, a revoked
key stays revoked on a stale store, and a krsk_ secret key presented as a
tenant key is refused. degraded and verified are the whole of what separates
this branch from an authentication bypass, so:
alert on
degraded, and onstats().policy.fetchedAtnot advancing;use
closed(orstale-then-closed) on any route where admitting an unknown key is unacceptable — money routes are the obvious ones;in a containerised deployment, reach for
stale-then-closedfirst. A fresh pod has no snapshot to be stale from, so it sits in the third row — every well-formed key admitted unverified — for the whole outage, and no volume configuration removes that for a pod that is born without a file. The cost is stated rather than softened: during an outage such a pod refuses every request with a 401 until its first successful poll. That is an availability cost against an authentication one, and it is the customer's call;stale-then-openremains the default.ship a disk snapshot (
persist, on by default) so a restart that keeps its filesystem is the second row rather than the third;poller.start()hydrates from it before its first poll for exactly this reason. This does not cover an immutable-container rollout, and mounting a volume is the second remedy there rather than the first: it narrows the window instead of closing it,emptyDirbuys nothing, a default rolling update and a scale-up are not covered by aReadWriteOncePVC or avolumeClaimTemplate, and a pod that does hydrate a pre-outage snapshot is the "Complete, stale" row — it serves, but revocations issued during the outage are not honoured. See@keyring-dev/cache's README, "Immutable containers and rollouts".The second row is the best case, not the only one. It assumes the hydrated file is one the node may believe. A file written before the SDK recorded its scope, or by a tenant-narrowed process, loads incomplete on purpose — the "Evicted …
complete: false" row — so until that node's first repair walk lands it verifies every key the file holds and admits any well-formed key that is not resident, including a forged one, where a node that (wrongly) believed the file would have answered 401. Bounded bymaxStalenessMs: past it a complete-but-stale store cannot prove a miss either.a node refuses a disk snapshot written under a different scope — a tenant-narrowed process must not hydrate a project-scoped process's file — and starts cold instead, reporting it once through
onError. Give each scope its owncacheDirwhen several processes share one filesystem.
verified: false also refuses idempotency outright, because there is no scope
to key a record on — see
A request we cannot scope.
That argument is load-bearing, so the code honours it rather than assuming it.
PolicyStore.status().complete is false whenever the store cannot prove a
miss — after an eviction under maxCachedKeys, when the snapshot was narrowed
to some tenants, or when the poller stopped a paged snapshot walk before its last
page — and every miss from such a store lands in the degraded branch instead of
the 401 one.
A project this node cannot cache
A node whose project holds more live keys than its maxCachedKeys cannot ever
build a complete cache, and an incomplete cache under the default above admits
every well-formed key, unverified, for the life of the process. Documenting that
is not the same as choosing it, so the SDK does two things in order.
First it raises maxCachedKeys by itself, up to maxAutoCachedKeys
(38,000 by default). Measured on the path the poller actually takes — a page
of snapshot JSON through JSON.parse and into the store — the cost depends on
the shape of your keys rather than on their number:
| shape | bytes/key | at the 38,000 ceiling | | ------------------------------ | --------- | --------------------- | | 1 scope, no rate limits | 519 | 18.8 MiB | | 3 scopes, a two-rule limit set | 855 | 31.0 MiB | | 8 scopes, a five-rule set | 1 271 | 46.0 MiB |
against ~8 MiB at the 10,000-key default on the middle row. A re-seed costs
about 1.4× more than the steady state while the page in flight is still
uncollected, and one PATCH /v1/projects/:id on your project's default rate
limits makes every node in your fleet re-seed at the same moment — so size the
container for the peak, not the steady state. The raise is reported through
onError as a PolicyCacheAutoRaised — a notice, not a failure; the poll it
came from succeeded — and it is scrapeable:
keyring.stats().policyCache;
// { autoRaisedTo: 38000, maxAutoCachedKeys: 38000,
// snapshotPages: 4, snapshotTruncated: false, scopeTooLarge: false,
// diskSnapshotRefused: false, authFailed: false, authFailedRefusingAll: false }
// autoRaisedTo is null on a node that never had to grow.
// diskSnapshotRefused is true on a node that found a disk snapshot for its
// project and environment and refused it for being another scope's -- the
// difference between "no file" and "a file this node must not read".
// stats().cache.configuredMaxCachedKeys is what you asked for; .maxCachedKeys
// is what is in force now.If the project still does not fit, you get a PolicyScopeTooLargeError, and
what it does depends on whether this node has ever served:
| the node | what happens |
| ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| has never served (cold) | Refuses to start. The store is never loaded, the poller stops, and the SDK rethrows out of a microtask — an uncaught exception, so the deploy dies |
| is already serving (warm) | Keeps serving and keeps polling. Records the condition, reports it through onError on every poll, and sets stats().policyCache.scopeTooLarge |
The asymmetry is deliberate and it is the whole point. "Refuse to start" and "kill a process that is answering requests" are different actions: your API runs in the same process, and a floor raise re-seeds a whole fleet in the same second, so throwing at a warm node is a fleet-wide outage of your product from a routine action on ours. Stopping its poller would be worse still — a node that stops polling never learns another revocation.
Drain a warm node rather than letting the SDK kill it. scopeTooLarge is a
readiness predicate:
app.get('/readyz', (_req, res) => {
const { scopeTooLarge } = keyring.stats().policyCache ?? {};
res.status(scopeTooLarge === true ? 503 : 200).end();
});Your orchestrator then stops routing to it and fails the deploy, which is what "refuse to start" means to an SRE, without killing a process that is still answering. The condition clears itself when a later walk finds the project fits again, so raising the bound recovers the node on its next re-seed.
The fix, in the order the message names it. Raise maxAutoCachedKeys (or
maxCachedKeys, which the ceiling never undercuts) and pay ~855 bytes of heap
per cached key. Setting maxAutoCachedKeys equal to maxCachedKeys turns the
automatic step off and goes straight to the refusal, which is the right shape
for a node with a hard memory bound.
new Keyring({ maxAutoCachedKeys: 100_000 }); // ~82 MiB on the middle row abovetenantIds is not that fix, and the error says so in the same sentence that
offers it. Narrowing lets the node start, and a tenant-narrowed store is
complete: false by construction — see the table above — so every key
outside the filter is admitted unverified under the stale-then-open default,
permanently. That is the state the refusal exists to prevent, reached by
following a remedy. Narrow only with stale-then-closed or closed on the
routes where admitting an unknown key is unacceptable.
onScopeTooLarge takes the cold decision back — to page, to fail a readiness
probe, to exit with your own code. It is not called on a warm node. A no-op
there leaves the node serving on the store it has, which at boot is an empty
one: every well-formed key admitted unverified, with no successful poll coming.
When it fires. On a walk — a cold start, a restart with no readable disk
snapshot, or a delta the feed answers with a snapshot — and on every other way
a project-scoped store loses completeness, because that is the state, not the
walk. A delta that admits a key past the bound (ordinary growth: a project
crosses maxCachedKeys by minting keys, with nothing re-seeding), a disk
snapshot an earlier process wrote short and start() hydrates from before its
first poll, and a last page larger than the max_keys this client asked for all
re-seed the same way. persist is on by default, so for a restart that keeps
its filesystem the disk path is the ordinary restart, and it is a delta
rather than a walk — a fresh pod in an immutable-container rollout has no file
to hydrate and re-seeds with a full walk instead, same as any other cold node.
Two cases are deliberately not touched. A tenant-narrowed node that still
does not fit stays a truncation: its store is complete: false by construction
already, so refusing to serve would buy nothing the tenant filter had not already
cost. And a walk cut short by the page bound (MAX_SNAPSHOT_PAGES) stays a
truncation too: that bound is about a control plane answering a cursor with
almost no keys, and a bug on our side must not be able to stop a customer's fleet.
The sharp edge: a route that declares scopes is not scope-checked in
the fail-open branch, because checking against a fabricated empty scope set
would be inventing an answer rather than admitting we have none. Use
stale-then-closed (or closed) on a route where that is unacceptable:
new Keyring({
routes: {
'POST /v1/payments': { onUnavailable: 'closed' }, // money: refuse rather than guess
'GET /v1/health': { skip: true },
'/v1/internal/*': { onUnavailable: 'stale-then-closed' },
},
});Route keys are METHOD /path, /path or /prefix/*; the most specific rule
wins regardless of declaration order, and keyring.stats().unmatchedRoutes
names any rule that has never matched — a typo in a route key is otherwise a
silent downgrade to the fail-open default.
Rate limiting
Limits are policy, and policy already has a distribution channel. A key's effective limits — its own override, or the project default, resolved by the control plane so the layering is in one place and not in three SDKs — arrive on the same snapshot and delta feed the keys do.
That is what keeps the round trip off most requests: a key with no limits
never talks to us at all. handle() stays synchronous; the check is a
separate awaited step the adapter makes only when a limit actually applies.
The counters are central, in our Redis, reached over the same krsk_
credential and the same base URL the policy poller already uses — never by
handing your process Redis credentials. Every limit in a check is evaluated in
one round trip, and all the counters are read before any of them is charged, so
a caller over their daily quota does not keep burning their per-second one.
Two algorithms, and the choice is about semantics rather than speed (all four candidates measured within 11 % of each other):
| Limit shape | algorithm | Why |
| --------------------------------- | ----------- | ------------------------------------------------------------------------------------------ |
| N per second or minute | sliding | Current window plus a weighted slice of the previous one. No boundary burst. |
| N per day or per billing period | fixed | What a customer means by a quota. A boundary burst at day granularity is nobody's problem. |
scope: 'tenant' shares a counter across every key that tenant holds, which is
what "this organisation gets 10,000 a day" means and what survives a rotation.
scope: 'key' is per key.
Headers
Both spellings ship, and the legacy triple is on by default because every client library in production parses it.
RateLimit-Policy: "per-key-second";q=10;w=1, "per-tenant-day";q=10000;w=86400
RateLimit: "per-key-second";r=7;t=1, "per-tenant-day";r=9000;t=3600
RateLimit-Limit: 10 # the most constrained policy, not the first
RateLimit-Remaining: 7
RateLimit-Reset: 1
Retry-After: 1 # on a 429, never before the window rolls
Keyring-Rate-Limited-Policy: per-key-secondSet legacyRateLimitHeaders: false to emit only the current draft.
When the counters are unreachable
rateLimitOnUnavailable — open by default.
| Mode | What happens |
| -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| open | Served, degraded set, no RateLimit header at all — we do not know the numbers, and a fabricated remaining is a number your caller will act on. |
| closed | 503 with Retry-After. For a route where admitting unmetered traffic is the expensive failure: an LLM call, an outbound SMS, anything you pay per unit for. |
There is no stale-then-open here, because there is no stale state to serve
from: a counter is a shared number, not a cached record. A local one is a
different limiter — a static split, measured at +12 % overshoot
and −6 % under-delivery — and v1 does not ship it.
The wait is bounded (rateLimitTimeoutMs, 500 ms). Our availability is not
yours, and an unbounded wait on our socket is the one line that would make that
false.
Idempotency
Stripe's semantics, because every backend engineer already knows them: header
Idempotency-Key, up to 255 characters; the status and body of the first
request are stored whether it succeeded or not, so a replay of a request that
500'd returns the same 500; a reused key carrying a different request is a
422; a duplicate arriving while the first is still running is a 409 with
Retry-After rather than a block.
Two deliberate deviations, and one more below:
PATCHandPUTare covered too. "PUT is idempotent by definition" is true of the resource state and false of the email it sends.- TTL is configurable, 24 h by default and 7 days at most, because some clients retry from a dead-letter queue the next business day.
Scope is per (project, env, tenant), not per key. Per-key scoping breaks
rotation-with-overlap: the old key writes the record, the retry arrives on the
new key, and your customer is charged twice by the feature that exists to
prevent exactly that.
Mounting it
The fingerprint covers the method, path, sorted query and body — never the
headers, because a retry legitimately carries a different User-Agent or trace
header. It therefore needs a parsed body, which is why it is a second Express
mount, after your body parser:
const mw = keyring({ secretKey, projectId });
app.use(mw); // verify + rate limit, no body needed
app.use(express.json());
app.use(keyringIdempotency(mw.keyring)); // idempotency, after the parserFastify needs no second registration: the plugin hooks onRequest for
verification and the limit, and preHandler for idempotency.
The body is canonicalised before hashing — object keys sorted, arrays left alone — so a retry serialised by a different client library is the same request.
At-least-once, and we cannot fix it
If your process dies after your handler has committed but before the record is
written, the retry re-executes. The record is left in_flight, the retry finds
an expired lock, steals it and runs your handler again. This is genuine
at-least-once behaviour, it is a property of every middleware-shaped idempotency
layer, and the only real fix is a transaction spanning our store and your
database, which we do not have.
The mitigation is the standard one: make your handler's own write idempotent — a unique constraint on your side, keyed on something derived from the request — and pass the same key through. A process that dies while the response is streaming is fine: the record is written before the body is flushed.
A legitimately slow handler is fine too. The 30 s lease is refreshed by a heartbeat while your handler runs, so a 45 s handler is not double-executed by an eager retry.
Size, and when the store is unreachable
Responses over 256 KB are not stored. The record is kept, so conflict
detection still works, and a replay answers 409 response_too_large_to_replay
telling the client to check the resource — never a wrong body.
idempotency.mode — closed by default, which is the opposite of the rate
limiter's, and the asymmetry is the decision. A rate limit not enforced during
an outage admits excess traffic; an idempotency record not written executes a
payment twice, and the request in front of us has explicitly asked for
exactly-once by carrying the header. Refusing it with 503 is declining to
promise what we cannot deliver, to a caller that is holding a retry loop and
will ask again. A request with no Idempotency-Key never reaches this code, so
the blast radius is only the mutating requests whose clients were built to retry.
mode: 'open' serves them anyway, for a customer whose handlers are already
idempotent on their own side. idempotency: false removes the feature.
A 5xx your transient predicate accepts — by default Retry-After present, or
status 503 — releases the record instead of storing it, so the retry
executes. Their 503 is usually "my database was failing over", not "this
operation was attempted".
Only a 5xx. The predicate is consulted for status >= 500 and for nothing
else, whatever it returns, because a 2xx or a 4xx is an attempt that happened
and releasing its record re-executes the operation. 202 Accepted +
Retry-After + Location is the standard async-job convention, and a job
started by a POST is exactly the mutating retried request this feature exists
for — without the gate the retry started a second one.
A request we cannot scope
A request whose decision was not verified is refused idempotency.
req.keyring.verified is false on the stale-then-open branch — before this
node's first snapshot lands, after an eviction, past maxStalenessMs — and on
that branch we know the environment (it is inside the key's own hashed material)
and neither the project nor the tenant. There is no (project, env, tenant)
scope to key a record on, so closed answers 503 and open serves the
request with no exactly-once promise, exactly as each mode does everywhere else.
It is the same judgement closed already encodes: we do not promise
exactly-once when we cannot deliver it. Scoping the record on something
key-specific instead would be worse — the retry that arrives once the snapshot
has landed is verified, addresses the real scope, finds nothing, and re-executes
having been told it would not. keyring.stats().idempotency.unverified counts
these separately from unavailable: our store was reachable, the policy was
not.
On-demand fill
A miss the store cannot prove schedules an asynchronous policy refresh. The request that triggered it never waits: a "miss → single remote verify" design was deliberately rejected, because that branch is exactly what puts a round trip on a cold cache. The fill only makes the next request better.
It is bounded three ways — one fetch in flight, a floor between fetches, a ceiling per minute — because a scan of invented keys must not become a scan of our own control plane.
A node configured with tenantIds is the one case a refresh cannot repair: no
delta or snapshot on that filter will ever carry a tenant the filter excludes,
and deriving the tenant from an unknown key hash would mean asking us on the hot
path. Such a node calls keyring.admitTenants([...]) — the caller knows which
of their tenants is calling; the SDK does not.
The ids are Keyring tenant ids (the UUIDs the control plane minted), not
your own external_id. An id that is not one is ignored, and every id a refused
snapshot carried is rolled back out of the filter — the node keeps serving from
the scope it already had. Without both, a single typo in a vendor's
externalId → keyringTenantId mapping left the node's snapshot 404ing forever:
stale past maxStalenessMs, admitting every miss under the fail-open default,
and receiving no revocation. stats().policyPollFailures is the number to alert
on.
A refusal names the filter, though, not an id — a 404 (or a 400: the node
treats them identically) is equally what a rolling deploy, a proxy or a gateway
answers — so the rollback is a suspicion and not a verdict. One refusal holds
every id that request carried, out of the filter and against admitTenants,
for a cooldown. The node re-seeds on the scope that worked and then re-admits
the suspects by bisection: half of a refused set goes back in, a poll that
serves it acquits that half outright, and the search narrows to the other.
What two refusals buy is the opposite of a hold: an id refused twice, at
least once on a request that named it alone, is convicted and leaves the
cooldown's protection for good — the node stops offering it, and only the
customer's own admitTenants or a restart brings it back. One refusal is not
enough for that, even alone, because the rolling deploy above answers the same
404: a recurring one that landed on a probe of a single id used to convict the
legitimate tenant it named. Two of them still can. If your tenant list is
machine-generated and your control plane 404s intermittently, a legitimate
tenant can be convicted and stay out until the process restarts; on a node with
traffic for that tenant the next cache miss restores it, on a quiet one nothing
does.
This matters because a tenant-narrowed store is not complete, so while a
tenant is out of the filter every request it makes is admitted unverified under
the fail-open default (or refused outright under closed): time out of the
filter is the cost, and re-admitting one id per poll made it grow with the
size of the refused burst — 50 tenants admitted alongside one bad id waited 50
polls.
There are two windows here and they are not the same length. The one that
matters to you is how long the tenant's keys are not in the store, because
that is how long its traffic is admitted unverified under the fail-open default
— or refused outright under closed. Measured on a 50-tenant burst, on a node
with no traffic of its own: 50 s with one bad id in it, 420 s with four,
790 s with eight, 1 150 s with sixteen and 1 725 s with thirty-two — it grows
with the number of ids your mapping got wrong. On a node whose request path
re-offers the tenants it cannot verify it is 50 s, and then 85–115 s at every
size. That is the number this mechanism should be judged on and the one to size
onUnavailable against.
The shorter window is membership of the filter, and it is the mechanism's own figure rather than yours: no id the node takes out of the filter stays out for more than one cooldown plus the poll that notices it is due, which is 60 s plus up to 30 s at the defaults: 322 of 366 swept shapes land in the 61.5–69.9 s band, on this revision and on every one before it. When the cooldown runs out the node puts the id back regardless of what the search has concluded, which is exactly what it did before there was a search at all. The "plus the poll" is not a hedge: a poll following a failed one backs off to as much as six intervals. Being in the filter for the duration of one request is not the same as being verifiable, which is why the two numbers differ by up to twenty times on the bursts above — and why the customer-facing one leads.
Two cases fall outside even that bound and both are earned: past
MAX_HELD_TENANTS there is no record left to put back with, and the id is
merely admittable again — the same thing, one step weaker — and an id the
search convicts leaves the bound for good, so that the node stops offering
a filter the control plane demonstrably refuses once a minute for ever. A
conviction takes two refusals for exactly that reason.
Putting an id back in the filter and putting it back on the wire are two
different acts, and only the first is on that clock. A poll that follows a
failed poll never carries a widening the control plane has not served: it
fetches a delta instead, which carries no tenant_id at all and so cannot be
refused for the filter. So the store stays fresh and the search keeps running
while the widening waits for the one poll it takes to get a success behind it.
Without that split the two properties really are incompatible — a snapshot
carries the whole filter, so re-admitting on the clock alone meant a dozen
refused ids arriving on a dozen polls could put a refusal on every poll for ever,
and only a restart cleared it. On top of that
the search is best effort: one bad id inside a thousand is convicted in
10–25 polls rather than 1,024, and a burst carrying several is measured rather
than promised, because separating K good ids from d bad ones against a yes/no
answer costs at least log₂(C(K+d, d)) experiments and no implementation beats
that. stats().tenantsHeldOut is how many are out, counting the one the node
is currently probing; it is bounded, so a caller passing a fresh UUID per
request cannot grow it without limit — and once that bound binds, widening
pauses for a cooldown, all widening, including tenants that have never been
refused, because past the cap an evicted id and a new one are the same thing to
this node.
admitTenants returns false for three different reasons — the id is already
in the filter, this node was never narrowed, or the value is not a tenant id —
so the one you have to act on has its own signal: a value dropped for shape is
counted in stats().tenantIdsRejected and reported to onError. No poll fails
in that case, so nothing else moves.
Telemetry
Fixed 10,000-event ring buffer (~2.5 MB), drop oldest when full, count the
drops, warn once a minute. Flush every 1,000 ms or 1,000 events, gzipped,
fire-and-forget with a 2 s timeout; one retry with jitter, then drop — we are
not a log shipper. A 429 is never retried: the Retry-After is honoured,
because retrying into a rate limit is the storm the limit exists to survive.
Above maxEventsPerSecond (5,000) non-error events are sampled and the rate is
recorded in the batch, but every 4xx, 5xx and degraded event is kept — the
customer debugs errors, not successes.
On SIGTERM and beforeExit the buffer is flushed with a 3 s cap. Adding a
SIGTERM listener suppresses Node's default action, so the handler re-raises
the signal once the flush is done and only when nothing else was listening — a
telemetry flush that quietly made a customer's containers un-killable would be a
far worse bug than losing a batch.
The query string is stripped from every path before an event leaves the process: a vendor whose callers pass a token in a query parameter would otherwise have it shipped to us and stored in our request log.
Log redaction
The single most common way an API key product leaks keys is its own
customers' logs. redact(value) deep-clones any value and masks a Keyring
key (kr_/krsk_) or an embed token wherever it sits — a bare string, a
header, a nested object, an Error's message — plus any field whose name
looks sensitive (password, secret, authorization, ...), whatever that
field's shape:
import { redact } from '@keyring-dev/sdk';
logger.info(redact({ headers: req.headers, err }));Pre-built redactors wrap this for pino, winston and bunyan, because none of the three offer a pattern-based hook on their own:
import {
pinoRedact,
winstonRedact,
bunyanRedactStream,
} from '@keyring-dev/sdk';
const logger = pino({ hooks: pinoRedact() });pinoRedact() returns both logMethod (masks the arguments passed to a log
call) and streamWrite (masks the final serialized line, which is the only
hook pino runs after its own req/res serializers and after a child
logger's bound fields are merged in — pass both, as hooks: pinoRedact()
does, or pino-http's request line and a child logger's bindings go out
unmasked).
format.combine(format(winstonRedact())(), format.json());
bunyan.createLogger({
name: 'app',
stream: bunyanRedactStream(process.stdout),
});bunyanRedactStream sees every line bunyan writes, whatever field it is
under, and is the one that gives full coverage. bunyanRedactSerializers
covers only the field names it names (err, error, req, res, key) —
cheaper when other bunyan streams already exist and should not be
double-wrapped, but a key under a field name it does not list goes out
unmasked; use the stream wrapper unless that trade is one you want to make.
redactString(value) is the same scan over one string, for a log line built
by hand rather than through a logger's own call.
Configuration
rateLimitOnUnavailable (open), legacyRateLimitHeaders (true),
rateLimitTimeoutMs (500), idempotency (true, or an options object, or
false).
secretKey, projectId and baseUrl default to KEYRING_SECRET_KEY,
KEYRING_PROJECT_ID and KEYRING_BASE_URL. A kr_ tenant key passed as
secretKey is refused: it cannot authenticate, and sending it would be the
customer's own credential leaving their process by mistake.
