npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@keyring-dev/cache

v0.1.1

Published

The Keyring SDK policy cache: versioned snapshot, delta poller, bounded LRU and snapshot-to-disk. Keeps @keyring-dev/core's verify() fed with no network call on the hot path.

Readme

@keyring-dev/cache

The policy cache behind @keyring-dev/core's verify(). Keeps a versioned snapshot of a project's keys in memory, refreshes it from a delta feed every 5 seconds, and keeps a copy on disk so a restart of the same process/filesystem is not a cold start.

@keyring-dev/core is framework-free and I/O-free, and stays that way. The poller, the LRU and the disk snapshot do I/O, so they live here. MIT, like core; it ships into a customer's process, and @keyring-dev/pepper is not in its dependency graph and never will be.

import { verify, lookupHashKey } from '@keyring-dev/core';
import { PolicyPoller } from '@keyring-dev/cache';

const poller = new PolicyPoller({
  baseUrl: 'https://api.example.test',
  secretKey: process.env.KEYRING_SECRET_KEY!,
  projectId: process.env.KEYRING_PROJECT_ID!,
});
// Hydrates from the on-disk snapshot if there is one, then polls. Restarting
// the same container/filesystem during a Keyring outage is invisible to your
// traffic. A fresh pod in an immutable-container rollout is a genuine cold
// start instead — see "Snapshot to disk" below.
poller.start();

// No await. No network call. This is the point.
const result = verify(request.headers.authorization!, {
  store: poller.store,
  env: 'live',
  requiredScopes: ['read:orders'],
});

Memory

≈ 640 bytes × maxCachedKeys, from measurement:

| maxCachedKeys | heap | bytes/key | p99 verify | | --------------- | ------- | --------- | ---------- | | 1 000 | 0.6 MiB | 679 | 3.0 µs | | 10 000 | 6.1 MiB | 641 | 3.6 µs | | 100 000 | 56 MiB | 591 | 3.8 µs | | 1 000 000 | 555 MiB | 582 | 4.8 µs |

The default is 10 000, so about 6 MiB. Memory is the finding here, not speed: an SDK that grows unbounded inside someone else's process is how you get uninstalled, and a customer running eight Node workers on a 2 GB container cannot afford the bottom row.

Re-measured on the path the poller actually takes — a page of snapshot JSON through JSON.parse, toKeyPolicy and into the store, with rateLimits on every record. These are the figures the auto-raise ceiling below is priced from:

| shape | bytes/key | 38 000 keys | | ------------------ | --------- | ----------- | | 1 scope, no limits | 519 | 18.8 MiB | | 3 scopes, 2 rules | 855 | 31.0 MiB | | 8 scopes, 5 rules | 1 271 | 46.0 MiB |

A key's cost is mostly its scopes and rate_limits, so the shape matters more than the record. Build the fixture as JS objects rather than as JSON text and every key shares one 'read:orders' string, which under-reports by 11–36 %; JSON.parse allocates per occurrence and is what the client does.

Four things follow from that, and they are worth knowing before you tune it:

  • The cache is bounded and evicts. store.stats() reports evictions and evictionRate. A non-zero eviction rate means maxCachedKeys is below your working set — raise it. Narrowing with tenantIds also stops the evictions and does not make the store complete: a narrowed store never is.
  • maxCachedKeys is also the page size the poller asks for, and it is a bound the poller will raise by itself. A project that does not fit gets maxCachedKeys raised to maxAutoCachedKeys (38 000, ~31 MiB) and the store stays complete. A project that does not fit under that gets a PolicyScopeTooLargeError whose first named remedy is raising the bound and paying for it. What that error does depends on whether the node has served: a node that never has refuses to start, and one that is already serving records the condition, reports it on every poll, keeps polling, and keeps serving — poller.scopeTooLarge is the readiness predicate to drain it on. Neither is silent: poller.autoRaisedMaxCachedKeys and a PolicyCacheAutoRaised on onError say the bound moved. The old behaviour — stop, leave status().complete false for ever, admit every well-formed key under stale-then-open — is the state this replaces, and the check is not only on the walk: a delta that admits past the bound, a disk snapshot an earlier process wrote short, and a last page larger than max_keys asked for all reach it too, and all of them re-seed.
  • A delta backfills a key you do not hold only when you hold the whole project. With tenantIds set, a key minted for a tenant this node never serves is not admitted: it would evict a hot key for one this process will never see, and such a store already reports complete: false. Without tenantIds the store's scope is the project, so it admits the key — the alternative is claiming completeness while missing every key minted since the snapshot, which would make "mint a key, then use it" fail on that node until it restarts. The rule is the tenant filter, not complete. A project-scoped store whose walk stopped at maxCachedKeys reports complete: false and still admits — its scope is the project either way, and skipping there would freeze its resident set at the project's oldest maxCachedKeys keys for the life of the process.
  • Eviction makes the cache incomplete, and store.status().complete says so. That matters: the reason a "key not found" can safely mean "key does not exist" is that the cache holds every key in its scope. Once it does not, a miss proves nothing. Use isAuthoritativeMiss() from @keyring-dev/core rather than reading undefined as a denial.

The snapshot is paged

The control plane will not serialise an unbounded snapshot into one response: 50,000 keys measured at 13.4 MB raw and 494 ms, on a process every customer's fleet is polling. Above its ceiling it answers a page and a next_cursor, and PolicyPoller walks the cursor to the end before it loads anything into the store — the store is replaced in one step, never filled page by page, because verify() is running against it throughout.

Two things follow, and both are visible from outside:

  • The whole walk reports one version, the first page's. A key minted mid-walk leaves the node's watermark behind that change rather than ahead of it, and the delta feed re-delivers it on the next poll. A watermark ahead of a change a node never received is the one error in this protocol that never corrects itself.
  • A walk this SDK cut short is not complete. snapshotTruncated says so and status().complete is false for the same poll. isAuthoritativeMiss() reads that, so a miss from such a store proves nothing and the fail-open policy applies. It keeps taking deltas: a key minted after the walk is admitted like any other. Two paths still reach it, and neither is "the project is large" any more — that one raises the bound and then refuses. A tenant-narrowed walk that does not fit truncates, because such a store is complete: false by construction anyway. And a walk that runs out of MAX_SNAPSHOT_PAGES truncates, because that bound is about a control plane answering a cursor with almost no keys, and a bug on our side must not be able to stop a customer's fleet.

A client that does not send max_keys is refused with 413 rather than handed a first page — a truncated snapshot that still claims to be complete is how a valid key gets a 401. This SDK always sends it, so it is never refused; the rule is what makes the protocol safe for a client that has not implemented paging yet.

Freshness

PolicyStore carries the version and the age of what it holds, so a caller can tell "absent from a fresh, complete snapshot" (the key does not exist) from "I have no fresh snapshot" (we do not know):

import { isAuthoritativeMiss, stalenessMs } from '@keyring-dev/core';

const status = poller.store.status();
// { snapshotVersion, fetchedAt, source: 'none' | 'disk' | 'network', complete }

if (!result.ok && result.reason === 'unknown_key') {
  if (isAuthoritativeMiss(status)) {
    // The key does not exist. Denying is correct.
  } else {
    // We do not know. This is where the fail-open policy applies.
  }
}

get() and status() are both synchronous, on purpose: an async signature would make a network call on the hot path expressible, and the whole architecture rests on it not being.

Propagation

A revocation reaches a node on its next poll, so ≤ 5 s (p100), ~2.5 s expected at the default interval, against a healthy control plane. policyRefreshMs: 1000 buys ≤ 1 s at five times the poll cost.

The interval is jittered by up to 20 % so a fleet that deployed together does not poll in lockstep, and the jitter is subtracted, never added — a poll can be early, never late, so the bound above stays a bound. The next poll is scheduled from when the previous one started, not from when it finished, so the interval governs when a node asks.

It still has to hear the answer back, so the honest bound is max(interval, round trip) + round trip. At the default 5 s interval and an EU round trip that is ≈ 5 s; a control plane answering slowly rather than failing widens it, up to the 10 s requestTimeoutMs. The ≤ 5 s figure is a statement about a healthy control plane, not a guarantee under one that is degraded.

For a hard answer rather than an interval, revoke with POST /v1/keys/:id/revoke?wait=true: it blocks until every node polling that project has acknowledged a version at or above the revoking one, and tells you which nodes did.

Snapshot to disk

On by default, because a restart during a Keyring outage otherwise takes the customer down — the genuinely dangerous case. This protects a restart that keeps its filesystem; see "Immutable containers and rollouts" below for the deployment shape where it does not apply.

  • Location: cacheDir, default os.tmpdir()/keyring-policy-cache.
  • Permissions: the directory is created 0700, the file 0600, and the write is atomic (temp file, fsync, rename).
  • Written off your main thread, one at a time. The synchronous form was 34.4 ms of fsync and serialisation per write at the 10,000-key default, landing on whatever request happened to be in flight — 17× the whole measured middleware budget. The fsync stays, because it is what makes the write atomic; what moved is where the waiting happens. Writes that pile up behind one in flight collapse into a single follow-up that reads the store when it starts, so the file converges on the newest state rather than replaying every step. await keyring.close() (or poller.flushPersist()) waits for it before the process exits.
  • Read on start(): the poller hydrates from disk before its first poll whenever it has no snapshot yet, so a restart that keeps the filesystem is invisible during an outage without your code doing anything — see "Immutable containers and rollouts" below for the case where there is no file to read.
  • Verified on load: a corrupt, truncated, foreign-format, oversized or unsafely-located snapshot is ignored, not trusted, and never fatal. The digest is integrity and not authenticity, so every record is validated before it can reach verify().
  • Scoped on load: the file records whether the process that wrote it was project-scoped or narrowed to a tenant filter, and a node refuses one it cannot load under its own scope — it starts cold and reports it once through onError (PolicySnapshotScopeMismatch, also on stats().policyCache.diskSnapshotRefused). See "Immutable containers and rollouts" for which pairs load and why the rule is not symmetric.
  • Contents: lookup hashes (the unpeppered SHA-256 the SDK computes for itself anyway), key and tenant ids, display prefixes, scopes and expiries. Never a plaintext key. Never the pepper. Neither exists in this process.

os.tmpdir() is world-writable, so the directory is the exposure, not the file: if it already exists and is a symlink, is owned by another user, or is writable by group or other, the SDK refuses to write there and reports it through onError rather than putting your policy somewhere someone else controls — and refuses to read from it too, silently, because a snapshot someone else could have planted there would otherwise be loaded and trusted at boot. The file is opened O_NOFOLLOW and refused unless this process owns it and no one else can read it. If you share a host with untrusted local users, set cacheDir to a private path you own. Set persist: false to opt out entirely.

Immutable containers and rollouts

The disk snapshot survives a restart that keeps its filesystem: a crash-loop restart of the same container, a classic VM or systemd deploy, or a cacheDir on a mounted volume all keep it.

An immutable-container rollout does not. A new pod starts with an empty cacheDir, so there is no file to hydrate from and the node is a genuine cold start — not "stale", nothing to be stale from. Scale-up, a node change and a pod eviction are the same case: whatever creates the new filesystem is what erases the snapshot.

What that means during a Keyring outage, plainly: the customer's traffic is not taken down, but under the stale-then-open default every new pod admits any well-formed key unverified (the checksum is CRC-32, so forging one costs nothing) until the control plane comes back. This is an authentication gap for the life of the outage, not a latency or availability footnote.

It also means such a node is cold for the scope-too-large boot refusal (see "Memory" above): a project past the cache ceiling fails the whole rollout — every new pod refuses to start — rather than surviving as one restart would with a warm disk.

Two remedies, in this order. The order is the point: the first closes an authentication gap, the second closes an availability one, and for an authentication vendor those are not interchangeable.

  • onUnavailable: 'stale-then-closed' (or closed on the routes that matter). This is the remedy to reach for first in a containerised deployment, because it is the only one that closes the gap above rather than making it less likely. Its cost, plainly and not as an "availability footnote": during a Keyring outage, every request a freshly started pod receives is refused with a 401 until that pod's first successful poll — there is no snapshot to be stale from, so there is nothing to serve. A pod that did poll successfully keeps serving its cached policy however stale, so the refusal is scoped to pods born during the outage. The trade is a refused request against an unverified one, and it is the customer's call to make — stale-then-open stays the SDK default, and this is a recommendation about a deployment shape, not a change to it.
  • A cacheDir on a volume that actually survives the way your rollout replaces pods. Second, because it narrows the window rather than closing it, and because it is not one line for every shape — the obvious guess is a no-op. emptyDir is created empty per pod and deleted with it, so it buys nothing across a rollout. A ReadWriteOnce PVC cannot attach to the surge pod of a default rolling update, so that shape is not covered either. What does work is a StatefulSet's volumeClaimTemplate, which follows pod replacement at the same ordinal — a crash-loop or a StatefulSet rolling update — but not scale-up, where a new ordinal gets its own empty volume; or a host-backed or shared volume (hostPath, or an RWX volume), which covers scale-up too at the cost of reopening the directory-ownership question two paragraphs above: it is the same world-writable-directory trust boundary as os.tmpdir() and needs the same private-path discipline. Whichever type is used, the cost is real and is not free just because it "restores the documented behaviour": a pod hydrating a pre-outage snapshot is the "Complete, stale" row of @keyring-dev/sdk's fail-open table — it serves, but revocations issued during the outage are not honoured for as long as the outage lasts. That row is the best case, and it needs the hydrated file to be one this node may believe: a pre-upgrade or narrowed-writer file loads incomplete, which puts the pod in the "Evicted … complete: false" row — it still verifies every key the file holds, and admits any well-formed key that is not resident — until its first repair walk lands. Not the "never loaded" row: such a pod does have a snapshot and reports source: 'disk'. See "the way it is made safe is a refusal" below.

None of this needs pods to coordinate — convergence always comes from the delta feed on the poll interval, never from the disk. Multiple processes sharing one container's cacheDir — several workers in one container is the ordinary shape of that — is safe, and the way it is made safe is a refusal:

  • The snapshot filename encodes (baseUrl, projectId, env), so one file per scope per filesystem — a shared temp directory is not a listing of the projects a host serves. It does not encode tenantIds, so two processes on the same project and env with different tenantIds do share one file.

  • The file records the scope of the process that wrote it, and a node refuses one that is not loadable under its own. Without that record the loader took scopeComplete from the file and narrowed from its own options, and a tenant-narrowed process reading a project-scoped file reached complete: true with narrowed: true — a combination no network walk can produce. Such a store drops every non-resident delta upsert while still claiming a miss proves the key does not exist, so a freshly minted key for a tenant that process serves was denied, with no repair path: a narrowed node never re-seeds itself. Refused, the node starts cold instead — one poll's worth of cost — and says so once through onError, naming the scope that wrote the file and the scope reading it.

  • What "loadable" means is not symmetric, because the two directions do not cost the same thing. A project-scoped node loads any file for its project and environment, including one written by a narrowed process or by an SDK too old to record a scope at all: nothing such a file can contain is outside this node's own scope, so refusing would cold-start it for nothing. A tenant-narrowed node loads only a file written under the same set of tenant ids (order and duplicates do not matter). A wider writer is the defect above; a narrower or merely overlapping one is refused for the second half of the same defect — nothing re-walks a narrowed node, and a narrowed store skips non-resident delta upserts, so a tenant it serves and the file never held would be missing for the life of the process.

  • Loading a file is not believing what it says about its own completeness. scopeComplete is taken from the file only when the file's recorded scope is project; a narrowed writer's file, and one with no recorded scope, load incomplete — and on a project-scoped node the first poll re-seeds them with a full walk. (Nothing re-walks a narrowed node, as above; a narrowed store is complete: false by construction and has nothing to repair.) This is not belt and braces: a file is not written by a snapshot walk, it is written by whatever the store held, and a pre-check narrowed process that had hydrated a project-scoped file persisted scopeComplete: true while holding one tenant's keys. An upgraded project-scoped worker that believed such a file would report complete while missing every other tenant, never re-seed (nothing re-walks a store that says it is complete), answer those tenants' valid keys as proven misses — a 401 in every onUnavailable mode — and then persist the state as a first-class scope: {kind: 'project'} file that every later worker trusts and no diagnostic flags. The cost of not believing it, on a healthy control plane, is one snapshot walk, once, on the first poll.

    During an outage it costs more than a walk, and the cost runs the other way. Between hydrating such a file and its first repair walk, a node is incomplete rather than complete, so under the default stale-then-open a miss is admitted unverified — including a forged well-formed key, which the same node would have answered with a 401 while it (wrongly) believed the file. That window is bounded by maxStalenessMs (60 s by default): past it the two behave identically, because a complete-but-stale store cannot prove a miss either. It is the same trade the whole fail-open default makes, and stale-then-closed closes it — the honest incomplete state is the one this SDK will report, because complete about a store that is demonstrably missing keys is the one thing it may never say.

  • A filter that widened at runtime is the filter that gets recorded, since those are the tenants the keys were actually fetched under. So a process that booted with tenantIds: [a] and admitted b writes [a, b], and a restart of the original narrower configuration refuses its own file and starts cold. That is deliberate — loading it would put b's keys in a store configured not to serve them — and it is the one case where the refusal costs something a correct configuration did not ask for.

  • Give each scope its own cacheDir if you run differently-scoped processes on one filesystem. They otherwise take turns refusing and rewriting one file, which is correct but buys neither of them a warm start.

  • Upgrading is additive, and the window is one-directional. The scope is a new field in the same keyring.policy-snapshot.v1 format rather than a new version, because a version bump would cold-start every node in every fleet once. The price is that an older SDK sharing that cacheDir ignores the field entirely and keeps poisoning itself exactly as before until it is replaced. During a rolling upgrade of a shared-cacheDir container that window is real; it closes when the last old process is gone, and there is nothing to do about it beyond finishing the rollout.

  • The write is atomic (temp file, fsync, rename), so a reader never observes a torn file. It is single-flight only within one process: two processes writing the same path concurrently are last-write-wins in whichever order the renames land, which can leave the file older than what one of them had already persisted. Harmless in practice — the next poll takes a delta from wherever the file landed and catches up — but it is not a guarantee the file holds the newest write.

  • Convergence across pods comes from the delta feed on the poll interval, not from the disk — the snapshot only ever shortens a single process's own cold start.