@keyring-dev/cache
v0.1.1
Published
The Keyring SDK policy cache: versioned snapshot, delta poller, bounded LRU and snapshot-to-disk. Keeps @keyring-dev/core's verify() fed with no network call on the hot path.
Readme
@keyring-dev/cache
The policy cache behind @keyring-dev/core's verify(). Keeps a versioned snapshot
of a project's keys in memory, refreshes it from a delta feed every 5 seconds,
and keeps a copy on disk so a restart of the same process/filesystem is not a
cold start.
@keyring-dev/core is framework-free and I/O-free, and stays that way. The poller,
the LRU and the disk snapshot do I/O, so they live here. MIT, like core; it
ships into a customer's process, and @keyring-dev/pepper is not in its dependency
graph and never will be.
import { verify, lookupHashKey } from '@keyring-dev/core';
import { PolicyPoller } from '@keyring-dev/cache';
const poller = new PolicyPoller({
baseUrl: 'https://api.example.test',
secretKey: process.env.KEYRING_SECRET_KEY!,
projectId: process.env.KEYRING_PROJECT_ID!,
});
// Hydrates from the on-disk snapshot if there is one, then polls. Restarting
// the same container/filesystem during a Keyring outage is invisible to your
// traffic. A fresh pod in an immutable-container rollout is a genuine cold
// start instead — see "Snapshot to disk" below.
poller.start();
// No await. No network call. This is the point.
const result = verify(request.headers.authorization!, {
store: poller.store,
env: 'live',
requiredScopes: ['read:orders'],
});Memory
≈ 640 bytes × maxCachedKeys, from measurement:
| maxCachedKeys | heap | bytes/key | p99 verify |
| --------------- | ------- | --------- | ---------- |
| 1 000 | 0.6 MiB | 679 | 3.0 µs |
| 10 000 | 6.1 MiB | 641 | 3.6 µs |
| 100 000 | 56 MiB | 591 | 3.8 µs |
| 1 000 000 | 555 MiB | 582 | 4.8 µs |
The default is 10 000, so about 6 MiB. Memory is the finding here, not speed: an SDK that grows unbounded inside someone else's process is how you get uninstalled, and a customer running eight Node workers on a 2 GB container cannot afford the bottom row.
Re-measured on the path the poller actually takes — a page of snapshot JSON
through JSON.parse, toKeyPolicy and into the store, with rateLimits on
every record. These are the figures the auto-raise ceiling
below is priced from:
| shape | bytes/key | 38 000 keys | | ------------------ | --------- | ----------- | | 1 scope, no limits | 519 | 18.8 MiB | | 3 scopes, 2 rules | 855 | 31.0 MiB | | 8 scopes, 5 rules | 1 271 | 46.0 MiB |
A key's cost is mostly its scopes and rate_limits, so the shape matters more
than the record. Build the fixture as JS objects rather than as JSON text and
every key shares one 'read:orders' string, which under-reports by 11–36 %;
JSON.parse allocates per occurrence and is what the client does.
Four things follow from that, and they are worth knowing before you tune it:
- The cache is bounded and evicts.
store.stats()reportsevictionsandevictionRate. A non-zero eviction rate meansmaxCachedKeysis below your working set — raise it. Narrowing withtenantIdsalso stops the evictions and does not make the store complete: a narrowed store never is. maxCachedKeysis also the page size the poller asks for, and it is a bound the poller will raise by itself. A project that does not fit getsmaxCachedKeysraised tomaxAutoCachedKeys(38 000, ~31 MiB) and the store stays complete. A project that does not fit under that gets aPolicyScopeTooLargeErrorwhose first named remedy is raising the bound and paying for it. What that error does depends on whether the node has served: a node that never has refuses to start, and one that is already serving records the condition, reports it on every poll, keeps polling, and keeps serving —poller.scopeTooLargeis the readiness predicate to drain it on. Neither is silent:poller.autoRaisedMaxCachedKeysand aPolicyCacheAutoRaisedononErrorsay the bound moved. The old behaviour — stop, leavestatus().completefalse for ever, admit every well-formed key understale-then-open— is the state this replaces, and the check is not only on the walk: a delta that admits past the bound, a disk snapshot an earlier process wrote short, and a last page larger thanmax_keysasked for all reach it too, and all of them re-seed.- A delta backfills a key you do not hold only when you hold the whole
project. With
tenantIdsset, a key minted for a tenant this node never serves is not admitted: it would evict a hot key for one this process will never see, and such a store already reportscomplete: false. WithouttenantIdsthe store's scope is the project, so it admits the key — the alternative is claiming completeness while missing every key minted since the snapshot, which would make "mint a key, then use it" fail on that node until it restarts. The rule is the tenant filter, notcomplete. A project-scoped store whose walk stopped atmaxCachedKeysreportscomplete: falseand still admits — its scope is the project either way, and skipping there would freeze its resident set at the project's oldestmaxCachedKeyskeys for the life of the process. - Eviction makes the cache incomplete, and
store.status().completesays so. That matters: the reason a "key not found" can safely mean "key does not exist" is that the cache holds every key in its scope. Once it does not, a miss proves nothing. UseisAuthoritativeMiss()from@keyring-dev/corerather than readingundefinedas a denial.
The snapshot is paged
The control plane will not serialise an unbounded snapshot into one response:
50,000 keys measured at 13.4 MB raw and 494 ms, on a process every customer's
fleet is polling. Above its ceiling it answers a page and a next_cursor, and
PolicyPoller walks the cursor to the end before it loads anything into the
store — the store is replaced in one step, never filled page by page, because
verify() is running against it throughout.
Two things follow, and both are visible from outside:
- The whole walk reports one version, the first page's. A key minted mid-walk leaves the node's watermark behind that change rather than ahead of it, and the delta feed re-delivers it on the next poll. A watermark ahead of a change a node never received is the one error in this protocol that never corrects itself.
- A walk this SDK cut short is not
complete.snapshotTruncatedsays so andstatus().completeis false for the same poll.isAuthoritativeMiss()reads that, so a miss from such a store proves nothing and the fail-open policy applies. It keeps taking deltas: a key minted after the walk is admitted like any other. Two paths still reach it, and neither is "the project is large" any more — that one raises the bound and then refuses. A tenant-narrowed walk that does not fit truncates, because such a store iscomplete: falseby construction anyway. And a walk that runs out ofMAX_SNAPSHOT_PAGEStruncates, because that bound is about a control plane answering a cursor with almost no keys, and a bug on our side must not be able to stop a customer's fleet.
A client that does not send max_keys is refused with 413 rather than
handed a first page — a truncated snapshot that still claims to be complete is
how a valid key gets a 401. This SDK always sends it, so it is never refused;
the rule is what makes the protocol safe for a client that has not implemented
paging yet.
Freshness
PolicyStore carries the version and the age of what it holds, so a caller can
tell "absent from a fresh, complete snapshot" (the key does not exist) from
"I have no fresh snapshot" (we do not know):
import { isAuthoritativeMiss, stalenessMs } from '@keyring-dev/core';
const status = poller.store.status();
// { snapshotVersion, fetchedAt, source: 'none' | 'disk' | 'network', complete }
if (!result.ok && result.reason === 'unknown_key') {
if (isAuthoritativeMiss(status)) {
// The key does not exist. Denying is correct.
} else {
// We do not know. This is where the fail-open policy applies.
}
}get() and status() are both synchronous, on purpose: an async signature
would make a network call on the hot path expressible, and the whole
architecture rests on it not being.
Propagation
A revocation reaches a node on its next poll, so ≤ 5 s (p100), ~2.5 s
expected at the default interval, against a healthy control plane.
policyRefreshMs: 1000 buys ≤ 1 s at five times the poll cost.
The interval is jittered by up to 20 % so a fleet that deployed together does not poll in lockstep, and the jitter is subtracted, never added — a poll can be early, never late, so the bound above stays a bound. The next poll is scheduled from when the previous one started, not from when it finished, so the interval governs when a node asks.
It still has to hear the answer back, so the honest bound is
max(interval, round trip) + round trip. At the default 5 s interval and an
EU round trip that is ≈ 5 s; a control plane answering slowly rather than
failing widens it, up to the 10 s requestTimeoutMs. The ≤ 5 s figure is a
statement about a healthy control plane, not a guarantee under one that is
degraded.
For a hard answer rather than an interval, revoke with
POST /v1/keys/:id/revoke?wait=true: it blocks until every node polling that
project has acknowledged a version at or above the revoking one, and tells you
which nodes did.
Snapshot to disk
On by default, because a restart during a Keyring outage otherwise takes the customer down — the genuinely dangerous case. This protects a restart that keeps its filesystem; see "Immutable containers and rollouts" below for the deployment shape where it does not apply.
- Location:
cacheDir, defaultos.tmpdir()/keyring-policy-cache. - Permissions: the directory is created
0700, the file0600, and the write is atomic (temp file, fsync, rename). - Written off your main thread, one at a time. The synchronous form was
34.4 ms of
fsyncand serialisation per write at the 10,000-key default, landing on whatever request happened to be in flight — 17× the whole measured middleware budget. Thefsyncstays, because it is what makes the write atomic; what moved is where the waiting happens. Writes that pile up behind one in flight collapse into a single follow-up that reads the store when it starts, so the file converges on the newest state rather than replaying every step.await keyring.close()(orpoller.flushPersist()) waits for it before the process exits. - Read on
start(): the poller hydrates from disk before its first poll whenever it has no snapshot yet, so a restart that keeps the filesystem is invisible during an outage without your code doing anything — see "Immutable containers and rollouts" below for the case where there is no file to read. - Verified on load: a corrupt, truncated, foreign-format, oversized or
unsafely-located snapshot is ignored, not trusted, and never fatal. The
digest is integrity and not authenticity, so every record is validated
before it can reach
verify(). - Scoped on load: the file records whether the process that wrote it was
project-scoped or narrowed to a tenant filter, and a node refuses one it
cannot load under its own scope — it starts cold and reports it once through
onError(PolicySnapshotScopeMismatch, also onstats().policyCache.diskSnapshotRefused). See "Immutable containers and rollouts" for which pairs load and why the rule is not symmetric. - Contents: lookup hashes (the unpeppered SHA-256 the SDK computes for itself anyway), key and tenant ids, display prefixes, scopes and expiries. Never a plaintext key. Never the pepper. Neither exists in this process.
os.tmpdir() is world-writable, so the directory is the exposure, not the file:
if it already exists and is a symlink, is owned by another user, or is writable
by group or other, the SDK refuses to write there and reports it through
onError rather than putting your policy somewhere someone else controls — and
refuses to read from it too, silently, because a snapshot someone else could
have planted there would otherwise be loaded and trusted at boot. The file is
opened O_NOFOLLOW and refused unless this process owns it and no one else can
read it. If you share a host with untrusted local users, set cacheDir to a
private path you own. Set persist: false to opt out entirely.
Immutable containers and rollouts
The disk snapshot survives a restart that keeps its filesystem: a crash-loop
restart of the same container, a classic VM or systemd deploy, or a cacheDir
on a mounted volume all keep it.
An immutable-container rollout does not. A new pod starts with an empty
cacheDir, so there is no file to hydrate from and the node is a genuine cold
start — not "stale", nothing to be stale from. Scale-up, a node change and a
pod eviction are the same case: whatever creates the new filesystem is what
erases the snapshot.
What that means during a Keyring outage, plainly: the customer's traffic is
not taken down, but under the stale-then-open default every new pod admits
any well-formed key unverified (the checksum is CRC-32, so forging one
costs nothing) until the control plane comes back. This is an authentication
gap for the life of the outage, not a latency or availability footnote.
It also means such a node is cold for the scope-too-large boot refusal (see "Memory" above): a project past the cache ceiling fails the whole rollout — every new pod refuses to start — rather than surviving as one restart would with a warm disk.
Two remedies, in this order. The order is the point: the first closes an authentication gap, the second closes an availability one, and for an authentication vendor those are not interchangeable.
onUnavailable: 'stale-then-closed'(orclosedon the routes that matter). This is the remedy to reach for first in a containerised deployment, because it is the only one that closes the gap above rather than making it less likely. Its cost, plainly and not as an "availability footnote": during a Keyring outage, every request a freshly started pod receives is refused with a 401 until that pod's first successful poll — there is no snapshot to be stale from, so there is nothing to serve. A pod that did poll successfully keeps serving its cached policy however stale, so the refusal is scoped to pods born during the outage. The trade is a refused request against an unverified one, and it is the customer's call to make —stale-then-openstays the SDK default, and this is a recommendation about a deployment shape, not a change to it.- A
cacheDiron a volume that actually survives the way your rollout replaces pods. Second, because it narrows the window rather than closing it, and because it is not one line for every shape — the obvious guess is a no-op.emptyDiris created empty per pod and deleted with it, so it buys nothing across a rollout. AReadWriteOncePVC cannot attach to the surge pod of a default rolling update, so that shape is not covered either. What does work is aStatefulSet'svolumeClaimTemplate, which follows pod replacement at the same ordinal — a crash-loop or aStatefulSetrolling update — but not scale-up, where a new ordinal gets its own empty volume; or a host-backed or shared volume (hostPath, or an RWX volume), which covers scale-up too at the cost of reopening the directory-ownership question two paragraphs above: it is the same world-writable-directory trust boundary asos.tmpdir()and needs the same private-path discipline. Whichever type is used, the cost is real and is not free just because it "restores the documented behaviour": a pod hydrating a pre-outage snapshot is the "Complete, stale" row of@keyring-dev/sdk's fail-open table — it serves, but revocations issued during the outage are not honoured for as long as the outage lasts. That row is the best case, and it needs the hydrated file to be one this node may believe: a pre-upgrade or narrowed-writer file loads incomplete, which puts the pod in the "Evicted …complete: false" row — it still verifies every key the file holds, and admits any well-formed key that is not resident — until its first repair walk lands. Not the "never loaded" row: such a pod does have a snapshot and reportssource: 'disk'. See "the way it is made safe is a refusal" below.
None of this needs pods to coordinate — convergence always comes from the
delta feed on the poll interval, never from the disk. Multiple processes
sharing one container's cacheDir — several workers in one container is the
ordinary shape of that — is safe, and the way it is made safe is a refusal:
The snapshot filename encodes
(baseUrl, projectId, env), so one file per scope per filesystem — a shared temp directory is not a listing of the projects a host serves. It does not encodetenantIds, so two processes on the same project and env with differenttenantIdsdo share one file.The file records the scope of the process that wrote it, and a node refuses one that is not loadable under its own. Without that record the loader took
scopeCompletefrom the file andnarrowedfrom its own options, and a tenant-narrowed process reading a project-scoped file reachedcomplete: truewithnarrowed: true— a combination no network walk can produce. Such a store drops every non-resident delta upsert while still claiming a miss proves the key does not exist, so a freshly minted key for a tenant that process serves was denied, with no repair path: a narrowed node never re-seeds itself. Refused, the node starts cold instead — one poll's worth of cost — and says so once throughonError, naming the scope that wrote the file and the scope reading it.What "loadable" means is not symmetric, because the two directions do not cost the same thing. A project-scoped node loads any file for its project and environment, including one written by a narrowed process or by an SDK too old to record a scope at all: nothing such a file can contain is outside this node's own scope, so refusing would cold-start it for nothing. A tenant-narrowed node loads only a file written under the same set of tenant ids (order and duplicates do not matter). A wider writer is the defect above; a narrower or merely overlapping one is refused for the second half of the same defect — nothing re-walks a narrowed node, and a narrowed store skips non-resident delta upserts, so a tenant it serves and the file never held would be missing for the life of the process.
Loading a file is not believing what it says about its own completeness.
scopeCompleteis taken from the file only when the file's recorded scope isproject; a narrowed writer's file, and one with no recorded scope, load incomplete — and on a project-scoped node the first poll re-seeds them with a full walk. (Nothing re-walks a narrowed node, as above; a narrowed store iscomplete: falseby construction and has nothing to repair.) This is not belt and braces: a file is not written by a snapshot walk, it is written by whatever the store held, and a pre-check narrowed process that had hydrated a project-scoped file persistedscopeComplete: truewhile holding one tenant's keys. An upgraded project-scoped worker that believed such a file would reportcompletewhile missing every other tenant, never re-seed (nothing re-walks a store that says it is complete), answer those tenants' valid keys as proven misses — a 401 in everyonUnavailablemode — and then persist the state as a first-classscope: {kind: 'project'}file that every later worker trusts and no diagnostic flags. The cost of not believing it, on a healthy control plane, is one snapshot walk, once, on the first poll.During an outage it costs more than a walk, and the cost runs the other way. Between hydrating such a file and its first repair walk, a node is incomplete rather than complete, so under the default
stale-then-opena miss is admitted unverified — including a forged well-formed key, which the same node would have answered with a 401 while it (wrongly) believed the file. That window is bounded bymaxStalenessMs(60 s by default): past it the two behave identically, because a complete-but-stale store cannot prove a miss either. It is the same trade the whole fail-open default makes, andstale-then-closedcloses it — the honest incomplete state is the one this SDK will report, becausecompleteabout a store that is demonstrably missing keys is the one thing it may never say.A filter that widened at runtime is the filter that gets recorded, since those are the tenants the keys were actually fetched under. So a process that booted with
tenantIds: [a]and admittedbwrites[a, b], and a restart of the original narrower configuration refuses its own file and starts cold. That is deliberate — loading it would putb's keys in a store configured not to serve them — and it is the one case where the refusal costs something a correct configuration did not ask for.Give each scope its own
cacheDirif you run differently-scoped processes on one filesystem. They otherwise take turns refusing and rewriting one file, which is correct but buys neither of them a warm start.Upgrading is additive, and the window is one-directional. The scope is a new field in the same
keyring.policy-snapshot.v1format rather than a new version, because a version bump would cold-start every node in every fleet once. The price is that an older SDK sharing thatcacheDirignores the field entirely and keeps poisoning itself exactly as before until it is replaced. During a rolling upgrade of a shared-cacheDircontainer that window is real; it closes when the last old process is gone, and there is nothing to do about it beyond finishing the rollout.The write is atomic (temp file, fsync, rename), so a reader never observes a torn file. It is single-flight only within one process: two processes writing the same path concurrently are last-write-wins in whichever order the renames land, which can leave the file older than what one of them had already persisted. Harmless in practice — the next poll takes a delta from wherever the file landed and catches up — but it is not a guarantee the file holds the newest write.
Convergence across pods comes from the delta feed on the poll interval, not from the disk — the snapshot only ever shortens a single process's own cold start.
