@krassx/turbokv
v0.1.0
Published
Layered in-memory KV cache for Node.js clusters: per-process L1 with a byte budget over a shared-memory L2 the primary owns and workers map read-only.
Maintainers
Readme
turbokv
A layered in-memory KV cache for Node.js clusters.
L1 is a per-process JS Map with a byte budget. L2 is a shared-memory
arena the primary owns and workers map read-only — the fd is opened O_RDONLY,
so a worker cannot write the arena even deliberately. Worker writes travel to the
primary through per-worker shared-memory submission rings rather than the cluster
IPC channel.
const { TurboKV } = require('@krassx/turbokv');
// or: import { TurboKV } from '@krassx/turbokv';
// primary, before forking
const cache = TurboKV.open();
TurboKV.install(require('cluster'));
// anywhere
cache.set('user:42', 'ada', { ttlMs: 60_000 });
cache.get('user:42'); // 'ada'Why the rings
process.send was never bandwidth-limited — the raw channel carries 439 MB/s
under JSON. Its costs are that each send synchronously freezes the sending
worker's event loop while V8 serialises the batch (0.49ms p50, 1.15ms p99 for a
~525KB batch), and that it is the one channel the application shares for its own
messages. Measured against a control doing identical work that never reaches the
channel, cache traffic degraded an application's own IPC round trip from 0.69ms
to 13.59ms at p99. Rings replace that with a memcpy:
| | delivered | shed | worker loop p99 | |---|---|---|---| | cluster IPC | 551k writes/s | 41.7% | 5.70ms | | shared memory | 1180k writes/s | 0% | 4.03ms |
Storage modes
| mode | stores | keeps Date/Map/Set | cost |
|---|---|---|---|
| bytes (default) | primitives + binary, no codec | rejects them | 1504 ns/op |
| safe | JSON | silently degrades them | 1789 ns/op |
| direct | v8.serialize | yes | 3250 ns/op |
direct uses each runtime's own v8.serialize, whose format differs between
Node and Bun. That is not reachable in practice: a cluster is built from
processes of one runtime, and an arena never outlives the primary that created
it — create() unlinks any prior segment and starts empty. Worth knowing only
if you attach to an arena from outside its own cluster, which is not a
supported arrangement.
Runtime support
Requires Node 18 or newer (Node-API level 8).
| | Node | Bun | Deno | |---|---|---|---| | addon, cluster, both transports | yes | yes | yes | | CJS + ESM entry points | yes | yes | yes | | post-collection heap guard | yes | yes | yes |
Installing
Prebuilt binaries ship inside the npm tarball, so a normal install needs no compiler and no network fetch. The addon is Node-API, so one prebuild per platform serves every supported Node major:
| | x64 | arm64 | |---|---|---| | Linux (glibc) | yes | yes | | Linux (musl) | yes | — | | macOS | yes | yes | | Windows | yes | — |
Anything not listed falls back to compiling from source at install time, which needs a compiler and Python, exactly as before.
Bun: this used to need
"trustedDependencies": ["@krassx/turbokv"], because Bun blocks lifecycle scripts by default and the addon was therefore never compiled. With prebuilds it no longer does —bun addstill reports "Blocked 1 postinstall", and the package works anyway, because the binary is already there. On a platform with no prebuild the old caveat still applies.
Bun runs the entire suite green — every unit test, both transports, and the full
primary-death recovery sequence — at roughly 15% below Node's throughput. The
heap guard works on all three: it is driven by a FinalizationRegistry rather
than gc performance entries, which Bun and Deno accept but never emit.
The direct codec used to return zero-filled typed arrays on Deno. v8.deserialize
does not copy an ArrayBufferView out of its input — it returns a view over it —
and the buffer it was given was a slice of Node's shared 8KB pool, so the value's
correctness rested on the runtime deriving an address from a non-zero
byteOffset. Deno 2.8.3 adds that offset twice. Decoding into an unpooled,
exactly-sized buffer removes the dependency (and stops each cached typed array
pinning a pool slab on every runtime); see DESIGN.md decision 46.
Releasing
Tag-driven. git tag v0.1.0 && git push --tags runs the release workflow: it
builds a prebuild on each target, verifies every one of them loads and passes
the suite, refuses to continue unless all six arrived and the tag matches
package.json, then publishes.
Authentication is trusted publishing (npm OIDC) — the workflow proves its identity to npm directly, so there is no long-lived token in repository secrets to leak, steal or forget to rotate, and provenance is attached automatically.
One exception: a trusted publisher is configured on a package that already
exists, so the first version cannot use it — and cannot use a token either.
npm restricts a token that bypasses 2FA to staging a publish, and staging
cannot bring a package into being; a CI token attempting it gets
E_STAGE_REQUIRED. So the first version is published from a developer machine,
behind the interactive 2FA the registry asks for, carrying the binaries from a
green release run rather than whatever one machine can build.
scripts/npm-oidc.sh does that and the setup around it: it checks the
preconditions, logs in through npm's browser OAuth flow, assembles the six
prebuilds from the CI run for the commit being published, publishes, runs
npm trust github to attach the trusted publisher, verifies it, and offers to
delete the NPM_TOKEN secret. Run it from a terminal, or with --yes from
something that has none.
That first version ships without provenance: the attestation is signed from CI's OIDC identity, which a local publish has no way to present. Every release after it goes through the workflow and carries one.
Layout
index.js index.mjs index.d.ts entry points; consumers never see src/
binding.gyp addon build, at the package root
src/ turbokv.js the JS layer (L1, coherence, transports)
native.js single place the addon is resolved
binding.cc *.h the arena, submission rings, platform layer
vendor/ rapidhash, verbatim upstream
test/ *_test.js run.js the suite; `npm test` runs run.js, so does CI
*.cc standalone C++ tests (arena, rings)
tsan/ types/ sanitizer gate, TypeScript declaration tests
bench/ microbenchmarks and design experiments
loadtest/ sustained multi-worker load harness (Docker)
scripts/ build helpersOperational notes
- Sizing: on Linux the arena is backed by
/dev/shm. Docker defaults it to 64MB — pass--shm-sizeorcreate()fails with a message naming it. - Worker ids start at 1.
0is the primary and is rejected. setnever throws. An unusable key, value or type returnsfalsewith the reason inlastError.- Primary death: a worker detaches, keeps serving its warm L1, polls, and
recovers when a heartbeat advances — then flushes L1 and re-claims a ring.
It notices on its next operation, not on a timer, so an idle worker holds
its mapping until something touches the cache. On Windows that matters: a
replacement primary cannot create the name while any handle is open, so
retry the restart rather than assuming one attempt after
primaryStaleMswill take. POSIX unlinks first and hides the difference. - Routing cluster messages yourself:
TurboKV.install(cluster)is the easy path and does the whole job. If your application owns the primary'smessagehandler instead, it must pass turbokv's messages toTurboKV.applyBatch()— on both transports.'shm'moves the writes off the channel; it does not take a worker off it. AclearAll()'s wipe, the two halves of its cluster-wide L3 clear generation, and the conditional re-time a worker asks for when anl3write fails all still travel as cluster messages and are applied nowhere else. Route the ring doorbell but not these and a value L3 rejected stays in the shared arena with no expiry, for every process, until something overwrites it. CallTurboKV.releaseWorker()on'exit'and'disconnect'for the same reason. transport: 'ipc'and thel3adapter: with an adapter attached, prefer the default'shm'transport. The primary drains every worker's submission ring before it decides whether an L3 read may be promoted, so a worker's write is never overwritten by the pre-write value L3 was still serving. An IPC-transport worker's write sits in its outbox, or in the cluster channel, where the primary cannot reach it — so for the one hop until that batch is delivered the primary may promote over it. The worker's own value wins once the batch is applied; the window is bounded, not closed.bigintvalues ontransport: 'ipc':process.sendserialises a batch as a unit and refuses abigintunder the default JSON serialization, so one such value drops every other key's write in the same batch — all of which already returnedtrue. Fork withserialization: 'advanced', or stay on'shm', where values never touch the channel.stats.flushDroppedcounts it.- Watching for that:
stats.l3FailTtlUnappliedcounts caps proven not to have landed, andstats.l3FailTtlUnconfirmedcounts those whose outcome the worker could not establish. Watch both, and expect the second one. The proof needs the invalidation ring to still reach back to the moment the cap was taken; the ring holdsringCaprecords —min(65536, 4% of the arena / 16 bytes), rounded down to a power of two with a floor of 8192, so 32768 at the 16MB default and 8192 on a 2MB arena — on the order of 100ms of primary writes under load, against roughly 7 seconds from a cap's mark to its retirement at the defaultl3FailTtlMs. So on a busy boxUnappliedgoes quiet andUnconfirmedbecomes the normal bucket — an operator watching only the first would see nothing during exactly the outage this exists for. transport: 'ipc',minLevel: 2or3, and a lapped ring: a worker that falls a fullringCapbehind the invalidation ring has to flush what it is holding, and it must then decide whether its own submitted-but-unapplied writes have landed. On the default'shm'transport it knows: the submission ring is single-producer, so the primary's consumer index answers exactly, and a write that has not been applied keeps missing locally. On'ipc'nothing acknowledges a batch, so the question is unanswerable and the mark is dropped — and that worker can then read L2's previous value for a key it wrote atminLevel: 2or3and was told succeeded. It is not one read: it lasts until something else replaces or invalidates the key. Nothing has to be broken to reach it — a synchronous write burst on the primary turns no event loop, so its doorbell never fires and its 500ms backstop never runs, and ~8300 records is a full lap on a 2MB arena. Stay on'shm'if you useminLevelabove 1.- The submission segment is mutually trusted; the arena is not. Workers map the arena read-only, which is the isolation that matters, but they map the submission segment read-write — one ring each, in one shared mapping. A buggy or hostile worker could always scribble anywhere in it and, because the primary snapshots its geometry once and bounds-checks every record, the worst it could do was lose its own writes. That is no longer quite the limit: the wrapped-ring fix above has each worker read its ring's consumer index to decide whether its write landed, so a worker that forges another worker's index can make that worker release a mark early and read a superseded value once, after a lap. The blast radius went from one worker's writes to one worker's reads. There is no trusted channel to check it against, and the trade buys a real stale read closed on the default transport for every correctly behaving process. If you run untrusted code in a worker, it does not belong in this cluster.
The full design, decision log and measurements — including what was built and rejected — are in DESIGN.md.
License
MIT. Vendors rapidhash (MIT).
