npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@sofa-buffers/corelib

v0.11.0

Published

Streaming, dependency-free TypeScript implementation of the SofaBuffers binary serialization format — usable from Node.js, browsers, Electron, and classic <script>.

Readme

SofaBuffers

Structured Objects For Anyone ... so optimized, feels amazing.

Would you like to know more?

SofaBuffers TypeScript library

CI Coverage Branches Docs

GitHub repository

A dependency-free, streaming TypeScript implementation of the SofaBuffers (Sofab) serialization format — the runtime stream core that runs anywhere JavaScript does (Node.js, browsers, Electron, Deno, Bun, a <script> tag).

Like protobuf's CodedInputStream / CodedOutputStream, it is meant to be driven by generated code: the sofabgen generator emits one class per message with serialize / decode methods that call these primitives. Decoding has one surface, the visitor (CORELIB_PLAN §5.3.1): a resumable push decoder that takes chunks of any size and calls one method per field.

Requirements

Node.js 20+ — CI runs 20 / 22 / 24 / 26 — or any modern browser / Electron / Deno / Bun. Built with TypeScript 6.x; targets ES2020 (bigint required).

Dependencies

None. Zero runtime dependencies; uses only standard JS / Web APIs (Uint8Array, DataView, TextEncoder / TextDecoder).

Feature flags

None — the build always ships every wire type.

Packaging

Published as @sofa-buffers/corelib:

npm install @sofa-buffers/corelib

Ships ESM (.js), CommonJS (.cjs), a browser IIFE global (SofaBuffers) and full type declarations.

Why this design

| Goal | How | |------|-----| | Runs everywhere | Pure TypeScript over Uint8Array / DataView / TextEncoder, no Node built-ins on the hot path. | | Streaming out | OStream writes into a small caller buffer and calls a FlushSink when it fills, so a message can exceed the buffer — by any amount, down to a one-byte buffer: a value too large for the buffer is split across flushes. | | Streaming in | IStream is a resumable state machine fed arbitrary chunks; large string / blob payloads arrive in pieces. | | One decode surface | The visitor, and nothing beside it (§5.3.1). decode() is that decoder fed once, so a whole-buffer decode runs the same code and the same rules as a chunked one. | | Full 64-bit fidelity | Scalars round-trip the entire uint64 / int64 range: number when exact, bigint beyond 2^53-1, and every integer callback carries the exact lo / hi halves beside the value for a bigint-free consumer (Long). | | Generated-code friendly | One flat Visitor per message, all methods optional; nesting arrives as sequenceBegin / sequenceEnd events carrying id and depth, which generated code routes on. | | Reserve-offset | new OStream(buf, offset) leaves room at the front for a lower-layer header, saving a copy. The offset belongs to that installation and is consumed by the flush that hands the unit over; setBuffer(buf, offset) from inside the sink re-arms it, for header room in every packet. | | Caller-owned buffers | The encoder allocates no output buffer, grows none, and has no hook that could grow one for it: it writes into yours, and when it fills it flushes to your sink, which may install the next buffer. growingOStream() is that caller ready-made — a scratch buffer with a sink that accumulates the result. | | No payload storage in the codec | After construction the encoder and decoder allocate no storage a wire number sizes (§6.6) — no views, no scratch, no growable state — apart from the language-forced handles of §6.6.2, itemised under Memory handling. Verified two ways: no allocation primitive on a codec path, and a flat heap over a complete encode and decode. | | No views | Nothing the decoder hands over aliases anything it owns (§6.7). A payload arrives as a range of the chunk you fed, so what you keep, you copied. | | Explicit endianness | IEEE-754 values are read / written little-endian — bit-for-bit identical on every engine, big-endian hosts included. | | Pluggable acceleration | The encoder's bulk array paths run through a swappable Kernel, and that interface is the entire seam: setKernel(yourKernel). No accelerated backend exists today — the kernel is the pure-TypeScript one on every host unless you build and install your own (native addon or WASM); the library ships no loader for one. |

Usage

The codec has four use cases — serialize a message that fits in one buffer, serialize one too large for the buffer (streamed out in chunks), deserialize a whole message, and deserialize one arriving in chunks — plus the generated-code path that wraps them.

Problems are reported by throwing SofabError; the cause is on SofabError.code (ARGUMENT, BUFFER_FULL, INVALID_MSG, INCOMPLETE, LIMIT_EXCEEDED), and that is the whole set. A read whose declared type contradicts the field on the wire is not an error at all: the field is skipped like an unknown id, the destination is left untouched and the decode stays COMPLETE. INVALID_MSG is a message malformed regardless of what follows; INCOMPLETE means the bytes merely ended inside a field, and is reported by what feed() returns rather than thrown — there is no finish/finalize step. LIMIT_EXCEEDED is neither: it is a receiver-local policy rejection, a field larger than a cap you configured (see Receiver limits).

Serialize

OStream writes into the buffer you hand it: the library allocates no output buffer and never grows one it was given. Where the schema bounds the message, that is one buffer of MAX_SIZE bytes:

import { OStream } from "@sofa-buffers/corelib";

const os = new OStream(new Uint8Array(MAX_SIZE));  // your buffer, sized from the schema
os.writeUnsigned(1, 42);
os.writeSigned(2, -7);
os.writeString(3, "hi");
const bytes = os.bytes();          // Uint8Array view of the finished message

Where it does not — no maxlen / count to size from — growingOStream() owns a buffer, hands it to the encoder like any other caller and replaces it with a bigger one of its own as the message grows:

import { growingOStream } from "@sofa-buffers/corelib";

const os = growingOStream();   // the accumulator owns the buffer
os.writeUnsigned(1, 42);
os.writeSigned(2, -7);
os.writeString(3, "hi");
const bytes = os.bytes();          // the whole message; never throws BUFFER_FULL

Everything below is written against OStream, and every one of its write* methods works the same on the stream growingOStream() returns.

Every integer written — scalar or array element — is checked against the 64-bit value domains: unsigned 0 .. 2^64 - 1, signed -2^63 .. 2^63 - 1. Anything outside them, and any number that is not an integer at all (a fraction, NaN, ±Infinity), throws SofabError with code ARGUMENT rather than a bare RangeError; the encoder never reduces a value modulo 2^64 and never puts a wrapped one on the wire. That answer does not depend on how the encoder was constructed, nor on the installed Kernel, which carries the same obligation.

The byte-level writeFixlen(id, data, subtype) is checked the same way against the fixlen domain: subtypes 0x4–0x7 are reserved, and an fp32 / fp64 payload is exactly 4 / 8 bytes. Either mistake throws ARGUMENT before a byte is written. String and Blob still take any length up to FIXLEN_MAX (0x7fffffff); the typed writeFp32 / writeFp64 / writeString are correct by construction.

Serialize stream

Constructed over a caller-owned buffer with a FlushSink, OStream drains that small buffer whenever it fills, so the buffer never has to be message-sized:

import { OStream, type FlushSink } from "@sofa-buffers/corelib";

const out: number[] = [];
// The sink is handed the installed buffer and the region's bounds — never memory
// from anywhere else (§5.1.6), and never a view the encoder built (§6.6).
const sink: FlushSink = (buf, start, end) => {         // or socket / file / stream
  for (let i = start; i < end; i++) out.push(buf[i]!);
};
const os = new OStream(new Uint8Array(16), 0, sink);   // tiny 16-byte buffer
for (let i = 0; i < 1000; i++) os.writeUnsigned(i, BigInt(i));
os.flush();                                            // push the tail

Nested sequences

A nested message is a sequence: a fresh id scope between a begin header and the 0x07 end marker. A sequence-typed field whose value equals its declared default is omitted from the wire, so the encoder holds the begin header back until the sequence proves it has content — no buffering of the sub-message:

const os = growingOStream();
os.writeUnsigned(1, 42);
os.writeSequenceBeginLazy(2);   // a nested field...
os.writeSequenceEnd();          // ...that got no content: header and end both vanish
os.writeSequenceBeginLazy(3);
os.writeString(1, "hi");        // content — commits the held-back header first
os.writeSequenceEnd();
os.bytes();                     // 08 2a 1e 0a 12 68 69 07  (field 2 is not on the wire)

Which closer to use is decided statically, by the position in the schema, not by the value:

| position | closer | |---|---| | struct / union field, array-field wrapper | writeSequenceEnd() — drops a contentless frame | | wrapper-array element, or an array field differing from a non-empty declared default | writeSequenceEndKeep() — always emits begin + end |

An element keeps its frame because element presence is what carries a dynamic array's length (highest present id + 1): dropping an all-default element would shorten the array. writeSequenceEndKeep() is the safe choice when in doubt — a needless one costs only a non-canonical empty frame that a decoder normalizes away — and raw transcoding, replaying bytes rather than encoding a schema value, uses it throughout so the output reproduces the input frame for frame.

const os = growingOStream();
os.writeSequenceBeginLazy(4);   // the wrapper array
os.writeSequenceBeginLazy(0);   // element 0 — has content
os.writeUnsigned(0, 7);
os.writeSequenceEndKeep();
os.writeSequenceBeginLazy(1);   // element 1 — all-default, but still present
os.writeSequenceEndKeep();      // ...so its frame stays: the array has length 2
os.writeSequenceEnd();
os.bytes();                     // 26 06 00 07 07 0e 07 07

Decoding is unaffected by the distinction: an empty frame is valid input that the message layer normalizes to the default, and an absent sequence field is reconstructed from the schema default. Nesting is capped at MAX_DEPTH (255) on both sides, and the encoder holds headers back to that full depth. A held-back header is encoder state, never buffer content, so streaming through a small buffer produces the same bytes.

Deserialize

decode() walks a whole buffer and calls one optional Visitor method per field; a field whose callback you did not implement is skipped. The visitor is flat: one object receives the whole message, nested scopes included.

import { decode, type Visitor } from "@sofa-buffers/corelib";

class My implements Visitor {
  a = 0;
  b = 0;
  unsigned(id: number, v: number | bigint) { if (id === 1) this.a = Number(v); }
  signed(id: number, v: number | bigint)   { if (id === 2) this.b = Number(v); }
  // fp32(), fp64(), string(), blob(), arrayBegin(), sequenceBegin(), ... as needed
}

decode(bytes, new My());

There are exactly two things to do with a field — read it or skip it (§6.7.2) — and not implementing a callback is how you say the second.

fieldBegin(id, wire) is announced first for every field — right after the header varint, before the value and before the value's own header word (a fixlen length word, an array count word, a nested sequence's fields). It gives a reader the field stream in wire order without writing the eight value callbacks. The sequence-end marker gets none: it closes a scope rather than opening a field, and its id is discarded (§4.9).

Do not apply a schema bound from it. An element id past the declared count looks decidable from the id alone, and is not: that bound applies only to a field whose subtype has confirmed it is the declared one, so it belongs on fixlenBegin. A message ending inside the fixlen word is INCOMPLETE even when the id would violate the bound. Throwing from fieldBegin is still how you reject a field the header alone settles — an id you will not accept in any shape. Everything schema-shaped stays on the later, more informative hook: a fixlen subtype and a declared length on fixlenBegin, a declared element count on arrayBegin.

An array's elements arrive through arrayBulk(id, kind, count), and only there: return the destination they should be written into — a number[], an exact-width typed array (Uint16Array, Int32Array, …), a Long[], a pair of Uint32Array halves, a Float32Array / Float64Array, or the raw fp32 words — together with the schema's element bound as min/max 32-bit halves, and the decoder fills it directly. Return null (or declare no arrayBulk) and the elements are walked over without being decoded at all. There is no callback per element: one array is one call.

The exact-width destination (typed) is for a field whose declared width is a typed array's own — a u16 array into a Uint16Array, a u64 array into a BigUint64Array. It stores unboxed, it costs two bytes an element rather than a tagged slot, and it lets the encoder know the width too: writeUnsignedArray reads a Uint16Array element without the type and range guards a number[] element needs, and reserves three bytes an element rather than ten, which is what a caller-sized buffer has room for.

A 64-bit destination is filled through the two 32-bit halves of each element, written into a Uint32Array over the array's own buffer — b[i] = 5n would demand a bigint per element, and this builds none. The same view reads them back when such an array is encoded. A boolean array has a destination of its own, bool, one byte per element: §4.4 gives a boolean no width bound, so a wire value of 256 is true and must NOT mask to 0 — the decoder normalizes every non-zero to 1, which is also the only value an encoder may write back.

The element bound is still compared, and that is not a formality: a typed array masks on store (a[0] = 70000 in a Uint16Array is 4464), while an element outside the declared width is a malformed message. So the width is matched against the bound once, at the hand-off — a destination narrower than the bound is refused with Argument, never silently truncated — and the fill then compares exactly as values does.

const big: number[] = [];
const v: Visitor = {
  arrayBulk: (id) =>
    id === 6 ? { values: big, minLo: 0, minHi: 0, maxLo: 0xffff, maxHi: 0 } : null,
};

sequenceBegin(id, depth) opens a nested scope and sequenceEnd(id, depth) closes it — the same visitor receives the scope's fields, with their own ids and depth + 1. Route on the (id, depth) pair, which a schema fixes statically:

let inChild = false;
const v: Visitor = {
  sequenceBegin: (id, depth) => { if (id === 3 && depth === 1) inChild = true; },
  sequenceEnd:   (id, depth) => { if (id === 3 && depth === 1) inChild = false; },
  unsigned: (id, value) => { /* `inChild` says which scope this id belongs to */ },
};

Return false from sequenceBegin to decline the whole subtree: no callback of any kind fires inside it, a scope opened within it is never offered either, and no sequenceEnd arrives for what was declined.

A declined subtree is still parsed — a sequence is framed by markers rather than by a length, so its end has to be found — but nothing is decoded into existence for it, and no receiver cap fires inside it: the callback that would have compared one is never called. Format ceilings (ARRAY_MAX, FIXLEN_MAX, MAX_DEPTH, the varint bound) apply inside a declined subtree exactly as outside it.

Deserialize stream

IStream resumes across chunk boundaries: feed it whatever the transport hands you and read the outcome from what feed() returns — COMPLETE or INCOMPLETE for the bytes consumed so far, the third outcome being thrown rather than returned (below). There is no end / finalize step and no second way to ask. The visitor is bound at construction. String / blob payloads arrive in one or more pieces, each a range [start, end) of the chunk you fed, tagged with the field's total length and the piece's offset within it:

import { IStream, DecodeStatus, type Visitor } from "@sofa-buffers/corelib";

const visitor: Visitor = {
  blob(id, total, offset, src, start, end) {
    /* copy `src[start..end)` to `offset` of a `total`-byte destination of yours */
  },
};

const is = new IStream(visitor);
let status = DecodeStatus.Complete;               // zero bytes end on a boundary
for await (const chunk of source) {
  status = is.feed(chunk);                        // any async byte source
}
// feed() never throws for a merely incomplete decode and never promotes one to
// an error (MESSAGE_SPEC §7). The caller owns end-of-input.
if (status !== DecodeStatus.Complete) {
  // stream ended inside a field (INCOMPLETE) — wait for more bytes, or treat
  // the truncation as an error if this really was the end of input.
}

The chunk is borrowed only for the duration of feed (§6.0): once it returns you may reuse, overwrite or free it, and what you decoded is unaffected — the decoder retains nothing that points into it. Copy what you want to keep, during the call; PayloadAcc and decodeUtf8 are the ready-made way.

feed() is the only way to ask. There is no status() accessor and no end step: what a feed returns, or what it throws, is the whole answer, so you are never one call short of knowing where you stand and never holding two answers that could disagree. (This library shipped that disagreement once — a status() that answered COMPLETE for a message feed had already refused — which is why the second way to ask is gone rather than repaired.) If you want the outcome again without keeping it, feed an empty chunk: it consumes nothing and returns the same value.

INVALID is terminal, and is the outcome feed() never returns: it travels on the error channel, as a thrown INVALID_MSG. A stream that has thrown it is poisoned for good — every further feed re-throws it without consuming a byte or calling the visitor, so a refused stream can never hand back a status at all:

import { SofabError, SofabErrorCode } from "@sofa-buffers/corelib";

const is = new IStream(visitor);
try {
  for await (const chunk of source) is.feed(chunk);
} catch (e) {
  if ((e as SofabError).code !== SofabErrorCode.InvalidMsg) throw e;
  // The verdict is in hand: `code` says INVALID_MSG, and nothing further needs
  // asking. Feeding on would only raise the same error again.
}

A receiver-side cap (LIMIT_EXCEEDED, see Receiver limits) is not the INVALID outcome and is never folded into one: the bytes are well-formed and decode under a looser cap. It travels the same error channel under its own code, and is terminal in the same way — the code you catch is what tells the two apart.

64-bit values without bigint

The default 64-bit surface is number-first: a value that fits exactly comes back as a number, and only past 2^53-1 is a bigint materialised — so the runtime type of a u64 / i64 depends on the value. Long, a value carried as two unsigned 32-bit halves (.low / .high), is the fixed-type alternative on the encode side. It is representation-only: the wire is identical to the number | bigint path, byte for byte.

import { Long, growingOStream } from "@sofa-buffers/corelib";

const os = growingOStream();
os.writeUnsignedLong(1, Long.fromValue(2n ** 63n));         // scalar
os.writeSignedLong(2, Long.fromValue(-(2n ** 62n)));
os.writeUnsignedArrayLong(3, [1n, 2n].map(Long.fromValue)); // array

On the decode side there is no channel to switch on: every integer callback carries the exact 64 bits as two unsigned 32-bit halves, beside the number-first value, and an array hands its elements over as Longs or as raw halves. Read whichever you want — the halves cost nothing to pass and nothing to ignore, and a Long built from them never goes through bigint arithmetic:

import { ArrayKind, decode, Long, type Visitor } from "@sofa-buffers/corelib";

const longs: Long[] = [];
const v: Visitor = {
  unsigned(id, value, lo, hi) { const x = Long.fromBits(lo, hi); },
  signed(id, value, lo, hi)   { /* lo/hi are the decoded two's-complement halves */ },
  // One call per array, not per element: the decoder fills `longs` itself. A
  // destination has to match the element kind, so decline the ones it does not:
  // `null` costs nothing but the call.
  arrayBulk: (id, kind) =>
    kind === ArrayKind.Unsigned
      ? { longs, minLo: 0, minHi: 0, maxLo: 0xffffffff, maxHi: 0xffffffff }
      : null,
};
decode(bytes, v);

Narrowing back is exact: value.low for u8..u32, and value.low | 0 for i8..i32.

Code generator

sofabgen compiles a schema to one class per message with a serialize (chaining OStream writes) and two decode entry points: a static decode for a message already in one buffer, and a static decoder() bound to IStream for one arriving in chunks — the same generated type driven whole-buffer or incrementally. Both drive the same visitor, because there is only one decode surface (§5.3.1): what changes is the drive, not the reader. A hand-written stand-in of both halves, encoded and decoded each way:

import {
  OStream,
  growingOStream,
  IStream,
  DecodeStatus,
  decode,
  type FeedStatus,
  type FlushSink,
  type Visitor,
} from "@sofa-buffers/corelib";

// generated by: sofabgen --lang typescript
class Point {
  x = 0;
  y = 0;

  serialize(os: OStream): void {
    os.writeSigned(1, this.x);
    os.writeSigned(2, this.y);
  }

  static decode(bytes: Uint8Array): Point {
    return _decodeIntoPoint(bytes, new Point());
  }

  /** The streaming half: a reader bound to the corelib's resumable IStream. */
  static decoder(): PointDecoder {
    return new PointDecoder();
  }
}

// The decode-into step sits beside the class, not on it: CORELIB_PLAN §6.1.1
// closes the generated object's surface to encode / decode / try_decode /
// serialize / deserialize / decoder, and `decode_from` / `decode_into` are two of
// the spellings it names as forbidden. It stays module-private — reachable from
// the sibling classes that decode into one another, and from nowhere else.

// Decodes into `o`, so a re-opened sequence continues the scope an earlier
// opening populated (MESSAGE_SPEC §7.4).
function _decodeIntoPoint(bytes: Uint8Array, o: Point): Point {
  decode(bytes, new PointVisitor(o));
  return o;
}

// generated alongside it: the visitor that fills a Point — the library's only
// decode surface (§5.3.1), one callback per wire type instead of one `case` per id.
// A visitor *is* the decode-into step: it writes into the object it was handed.
class PointVisitor implements Visitor {
  private readonly out: Point;
  constructor(out: Point) { this.out = out; }

  signed(id: number, v: number | bigint): void {
    if (id === 1) this.out.x = Number(v);
    else if (id === 2) this.out.y = Number(v);
    // no branch for an unknown id — or for a field whose wire type is not
    // `signed` — so it is skipped and the decode stays COMPLETE (§7.3)
  }
}

// ...and the handle Point.decoder() returns: an IStream plus its destination
class PointDecoder {
  readonly message = new Point();
  private readonly is = new IStream(new PointVisitor(this.message));

  // The one place the answer is: what feed() returns, or what it throws. A
  // status() accessor here would be a second way to learn the same fact, which
  // is the second way it can be learned wrong — so the generated handle does not
  // grow one either.
  feed(chunk: Uint8Array): FeedStatus { return this.is.feed(chunk); }
}

const p = new Point(); p.x = 3; p.y = 4;

// one-shot: encode into memory, decode a whole buffer
const os = growingOStream(); p.serialize(os);
const wire = os.bytes().slice();
const got = Point.decode(wire);            // got.x === 3, got.y === 4

// streaming out: the same serialize(), over a 4-byte buffer with a sink. The sink
// is handed the installed buffer and the region's bounds — never memory from
// anywhere else — so it copies out what it wants to keep.
const parts: Uint8Array[] = [];
const sink: FlushSink = (buf, start, end) => { parts.push(buf.slice(start, end)); };
const so = new OStream(new Uint8Array(4), 0, sink);
p.serialize(so); so.flush();               // the same bytes, in pieces

// streaming in: feed those pieces — or any other chunking — to the decoder
const dec = Point.decoder();
let st: FeedStatus = DecodeStatus.Complete;     // zero bytes end on a boundary
for (const part of parts) st = dec.feed(part);

// COMPLETE says the bytes so far ended on a field boundary, not that the
// message is over — the caller's framing decides that, and a still-INCOMPLETE
// status once the input really has ended is truncation (§5.2.4).
const streamed = st === DecodeStatus.Complete ? dec.message : null;

A generated visitor takes the nested cases too: a nested message switches the router into the child's fields on sequenceBegin(id, depth), and a compact scalar array is filled straight into the destination the router hands over at arrayBegin / arrayBulk, so no part of the message is ever buffered whole. Nothing from a fed chunk is retained either — a string is decoded and a blob copied on the way into the destination — so a chunk is reusable the moment feed returns.

This example is compiled and executed by the test suite (test/helpers/readme-generator-example.ts), so it cannot drift from the API.

Memory handling

Who owns the bytes:

  • Encode (OStream). Every buffer the encoder writes into is caller-supplied: the library allocates none of its own and never grows or reallocates one it was handed — new OStream(buf, offset?, flush?) writes into buf and into nothing else. When it fills it calls the flush sink with that buffer and the region's bounds — (buffer, start, end), never memory from anywhere else and never a view the encoder built, since pass-through is forbidden (§5.1.6) — and continues; the region is valid for the duration of that call. With no sink it throws BUFFER_FULL. bytes() returns a view of what is in the buffer — with a sink, only the not-yet-flushed tail — so .slice() it if it must outlive the next write.

  • The offset belongs to the installation, not to the buffer. It reserves room at the front of the unit the buffer-set begins — the constructor or setBuffer — and handing that unit to the sink consumes it: a sink that returns without installing a buffer has copied, so the encoder keeps writing into the same buffer and resumes at 0, with the whole buffer usable from there. A sink that wants header room in every flushed unit — one framing header per packet — re-arms it by calling setBuffer(buf, offset) from inside the callback; passing the buffer it already has counts. A sink that takes the buffer (hands it to a transport, queues it, gives it to DMA) must install a replacement before returning. Either way the bytes are the same — only the unit sizes differ. reset() and bytes() follow the current installation, so after a flush they are relative to 0; on a sink-less stream, which can never flush, the reservation stands for the life of the encode.

  • Encode into memory (growingOStream()). The allocating half is the caller's role. growingOStream(initialCapacity?) is that caller ready-made: a scratch buffer installed with a sink that accumulates the result (§5.1.2). It never throws BUFFER_FULL, its bytes() is the whole message (a view — .slice() it if it must outlive the next write or growth), and reset() keeps the buffer it grew to, so a pooled encoder stops allocating. It is an ordinary streaming stream otherwise, so setBuffer works and means what it always means: the not-yet-flushed bytes are dropped and encoding continues into your buffer. Pass an initialCapacity when you know roughly how large the message is: a message built from many small fields grows by doubling, so 100 KB of them costs nine enlargements from the 256-byte default. A single large field does not — a bulk write tells the accumulator how much contiguous room it wants, so the buffer reaches that size in one step and the write keeps its bulk route.

    It reaches the encoder through setBuffer(buffer, offset, carried), the third argument being how many bytes of the message the replacement already holds before offset. That is what keeps bytes() meaning "the message" across an enlargement, and it is available to any caller that keeps a message in one growing store.

    Its storage is carved from a shared slab while it is small enough (up to 4 KiB of an 8 KiB slab). A carve is handed out once and never recycled, so no two encoders ever share bytes and no message can read another's; what it changes is lifetime — a retained bytes() view keeps its slab alive, so .slice() (already the advice for a view that outlives the next write) is also what releases it.

  • MIN_OUTPUT_BUFFER = 1. The smallest buffer this port accepts for streaming, exported from the package so a caller can size from it. It is 1 because the encoder splits every atomic unit — field header, fixlen word, element count, a scalar or array element varint, an fp32 / fp64 element — across a flush, so a message of any size encodes through a one-byte buffer and the bytes produced are identical at every size. It binds a buffer installed with a sink, at construction and at every mid-stream setBuffer: buf.length - offset must be at least MIN_OUTPUT_BUFFER, and a smaller window is rejected right there with ARGUMENT — never partway through a message — leaving the encoder on the buffer it already had. A buffer installed without a sink has no minimum: no flush can occur, so nothing can be split, and a two-byte message encodes into a two-byte buffer.

  • Decode (decode() / IStream). You own the bytes being parsed, and they must stay valid only for the duration of the feed (or decode) call. After it returns, reuse, overwrite or free them freely: nothing the decoder produced points into them.

  • No views. The decoder exposes no zero-copy view of a decoded value, no payload-position getter and no borrowed value (§6.7) — on the one-shot path exactly as on the streaming one, with no option that reinstates one. A string / blob payload is reported in pieces as (src, start, end), where src is the chunk you fed: the decoder builds no view over it and keeps no storage of its own, so whatever you want to keep, you copy out of memory you already own, during the call. Scalars are delivered by value. If any of this README ever describes a borrowed decoded value, either the README or the port is wrong.

  • No wire value decides an allocation in the codec. After construction the encoder and decoder allocate no payload storage (§6.6), and nothing at all except the itemised handles below: no per-message, per-field or per-chunk allocation, no growable state, and no accumulator for a payload that straddles a chunk — a decoder's whole memory is fixed-size state sized from this format's constants (a MAX_DEPTH scope stack, a partial varint, an 8-byte float landing zone). Constructing an OStream / IStream is the one allocating step, and decode() reuses one decoder across calls so a one-shot caller does not pay it per message. A bigint for an integer past 2^53 is not an exception: it is a value, not storage, and the lo / hi halves beside it are there for a consumer that would rather not have one. A Long written into a longs or typed bulk destination is the same kind of thing — a value, placed in storage you supplied.

  • The language-forced handles, itemised (§6.6.2). JavaScript will not let a codec place or take an IEEE-754 value at a byte offset, or copy a range of bytes, without building an object first: TypedArray.set — the only memcpy there is — takes a typed array as its source, and a float needs a DataView. These are all of them:

    | handle | where | how many | |---|---|---| | DataView over the output buffer | Kernel, bulk fp32 / fp64 arrays | one per bulk call, and only from 64 fp32 / 16 fp64 elements up | | DataView over the fed chunk | IStream, bulk float array reads | one per chunk, on the first run in it that clears the same thresholds | | subarray of the caller's payload | OStream.writeRaw, as set's source | one per copied piece, only when the payload does not fit the buffer | | Uint32Array over a 64-bit destination | IStream, an array handed a BigUint64Array / BigInt64Array | one per array | | Uint32Array over a 64-bit source | Kernel, writeUnsignedArray / writeSignedArray from one | one per bulk call | | Uint32Array + DataView over an fp32 source and the output | Kernel, writeFp32Array from a Float32Array | one pair per bulk call, at every length | | Uint32Array over an fp32 source | OStream, writeFp32Array from a Float32Array that does not fit the buffer | one per call, however many flushes it spans |

    Each addresses storage you supplied, each is sized by that storage and never by a number from the wire, and none of them leaves the codec. The 64-bit Uint32Array rows are what a 64-bit element costs instead of a bigint: its halves are already in hand, and a view over the caller's own array is how they are written without building one.

    A scalar float, a float array below the element threshold, and a float array fed in chunks too small to hold a long run all build no handle at all: they go through a shared 8-byte scratch word, which is fixed state. The one exception is the two fp32-source rows — an fp32 array whose source is a Float32Array already holds the wire words, and copying them is what keeps a signaling NaN intact (§4.6/§6.5), so those handles are built at any length rather than past a threshold. Streamed through a buffer too small for the whole array, the words still go out as words — through the Uint32Array alone, stored with shifts — so a small buffer gives the same bytes as a large one (§5.1.4). heap-free-codec.test.ts asserts every count in this table exactly, including the short runs that allocate nothing, the element one under the threshold, and the two-element Float32Array that builds the pair anyway. The thresholds and what they were derived from are on FP32_HANDLE_MIN / FP64_HANDLE_MIN in the API documentation.

  • The bulk array hand-off borrows your destination until arrayEnd. Visitor.arrayBulk hands the decoder the array, Long[] or typed array it should fill for one array field; the decoder writes into it ascending from index 0 and holds it from the hand-off until that array ends, which on a chunked decode spans several feed calls. It is dropped there, and on decode()'s pooled machine when the call returns. The object must stay the same one for the whole array. A typed destination must already hold count elements; a plain number[] / Long[] grows as it fills and is cut to the elements written when the array ends — including to zero for an array that is empty on the wire — so reusing one across arrays or messages never leaves the previous array's tail behind and its length after arrayEnd is exactly that array's element count. An element the target's bound rejects stops the fill: everything before it is written, it and everything after it are not. This is the only reference the decoder keeps into your storage between calls.

  • The static helper layer allocates, on your behalf. PayloadAcc, ElementSeq, FramedSeq, StringSeq, BlobSeq, decodeUtf8, elementsEqual, longElementsEqual and fp32RawBytes are the generated layer's code shipped here for reuse (ARCHITECTURE §8), not part of the codec: the codec never calls them, and they allocate the values they build.

  • String validity is checked where a string is materialized (§6.4.5). JavaScript strings are a Unicode type, so this port is always strict — but a string payload piece is raw wire bytes and is not validated (it may end mid-code-point), so whoever materializes one owns the check. decodeUtf8(bytes, start?, end?) is that check, exported for exactly this: it rejects malformed bytes as INVALID_MSG rather than as a platform TypeError. Rolling your own instead means new TextDecoder("utf-8", { fatal: true }) — the default TextDecoder silently substitutes U+FFFD, which the format forbids in either direction, and TextEncoder does the same to an unpaired surrogate where this encoder refuses it with ARGUMENT.

  • Reassembly is the caller's, with a helper. The codec holds no payload across feed calls. PayloadAcc.take(total, offset, src, start, end) joins the pieces — one accumulator per decoder, since only one payload is ever in flight — and returns storage of its own that aliases nothing, on the whole-payload path exactly as on the split one. StringSeq / BlobSeq collect the elements of a string / blob wrapper array; ElementSeq holds the index rules for any element kind (index bound, gap fill, last-write-wins) and FramedSeq is its twin for an element whose default is a fresh object — a struct, a union, a nested row — where one shared default would alias every gap of the array onto a single instance; elementsEqual (and longElementsEqual, for Long-backed 64-bit arrays, whose elements are object identities) is the array form of the omit-if-default test an encoder applies before writing a field; fp32RawBytes turns the 32-bit word Visitor.fp32 hands over back into the four wire bytes a generated message keeps beside an fp32 it cannot re-encode from a number (§6.5).

Receiver limits

This corelib holds none, by design. A field the schema leaves unbounded is still bounded by the receiver — CORELIB_PLAN §6.2.1 admits no unset state and no unlimited mode — but the numbers belong to generated code, which knows the schema and the deployment, and §6.2.1 is explicit that a codec

MUST NOT hold a limit of its own, MUST NOT supply a default for one it was not given, MUST NOT read an omitted argument as unlimited, and MUST NOT clamp to one

and that a format ceiling reached because no cap was stated

is the format's bound, not a receiver cap, and a port MUST NOT present it as one.

So decode(bytes, visitor) and new IStream(visitor) take no limits argument. Up to v0.10.0 they took a DecodeLimits whose absent members fell back to ARRAY_MAX / FIXLEN_MAX; that object is gone, along with the LIMIT_EXCEEDED rejections it raised against ceilings nobody had configured.

| the cap | who states it | who compares it | |---|---|---| | max_dyn_array_count on an array field | generated code | its own arrayBegin | | max_dyn_string_len / max_dyn_blob_len on a string / blob field | generated code | its own fixlenBegin | | the element index of a wrapper array | generated code | StringSeq / BlobSeq / ElementSeq / FramedSeq, from receiverCap | | the element byte length of a wrapper array | generated code | StringSeq / BlobSeq, from receiverElemMax |

§6.2.1 permits the comparison to run inside the corelib — "A corelib MAY take a limit as an argument and perform the check itself" — and the collectors do exactly that, for the one shape a visitor cannot see: a wrapper array's element length words go to the collector, never to the generated visitor. Every one of their bounds is a required constructor argument with no default, the schema halves included (UNBOUNDED, -1, is the explicit "the schema declared none"):

//              out, acc,          count, maxlen, name,   receiverCap, receiverElemMax
new StringSeq(  out, new PayloadAcc(), UNBOUNDED, UNBOUNDED, "tags", 65_536, 1 << 20);
new BlobSeq(    out, new PayloadAcc(), 8,         4096,      "parts", 65_536, 1 << 20);
new ElementSeq( out, defaultElem,      UNBOUNDED, "rows",   65_536);
//              out, make,             count,     name,     receiverCap
new FramedSeq(  out, () => new Elem(), 8,         "codes",  65_536);

Each pair is exclusive, never additive (§6.2.1: a cap "MUST NOT be applied to a field the schema already bounds"): where the schema declared count / maxlen that bound governs and its violation is INVALID_MSG, a statement about validity; where it declared none the receiver cap governs and its violation is LIMIT_EXCEEDED, a policy rejection on well-formed bytes.

A receiver bound that is about to govern and states no cap — negative, NaN (what an omitted argument becomes in JavaScript), Infinity — is refused at construction with SofabErrorCode.Argument. It is neither of the two categories above: no receiver policy was set, so there is no LIMIT_EXCEEDED to raise ("a format ceiling reached because no cap was stated is the format's bound … and a port MUST NOT present it as one"), and no unlimited mode to fall back to ("MUST NOT read an omitted argument as unlimited"). It is a mistake in the call, which §6.3 makes InvalidArgument. The refusal is fail-closed: nothing is decoded through a collector that could not be built. A bound the schema half makes inert is not checked — §6.2.1 forbids applying it at all.

What this decoder still owes the layer that holds the numbers is the enforcement point §6.2.1 requires — the count / length header, before the allocation the cap exists to prevent, and behind the MESSAGE_SPEC §7.3 tag test:

  • arrayBegin(id, kind, count) is raised at the count word, before any element is delivered (for a fixlen array, at its element-length word — the element kind is unknown until then, and still before any element);
  • fixlenBegin(id, subtype, total) is raised at the length word, before any payload piece;
  • both carry the number a destination gets sized from, so a rejection there costs no allocation at all. Reject, never clamp: materialising limit elements where the wire said more is data corruption wearing a safety jacket.

A cap therefore applies only to a field you read, and that falls out of the structure rather than needing a rule: a field the visitor steps over never reaches the callback that holds the number, so a decode that walks past an over-cap field it was never going to read stays COMPLETE (§6.2.1's "a skipped field is never capped"). The format ceilings are not yours to waive this way: a count above ARRAY_MAX or a length above FIXLEN_MAX stays INVALID whether anyone reads the field, because it bounds what the wire may express.

A cap rejection is raised by throwing SofabError with code SofabErrorCode.LimitExceeded, which propagates out of feed / decode. It is not the INVALID outcome and is never folded into one — the same bytes decode under a looser cap. The error channel is the only place it appears, and that is not a gap: the three-valued outcome has no value for "valid, but more than I am configured to accept", so there is nothing about it a returned status could have said. Catch it and read its code.

Build & test

npm ci
npm run typecheck      # tsc --noEmit (strict)
npm test               # vitest run: vectors, chunked feeding, memory rules, round-trips
npm run coverage       # vitest run --coverage (v8)
npm run build          # tsup -> ESM + CJS + IIFE + .d.ts in dist/
npm run smoke          # cross-runtime smoke test of the built bundle

Tests live in test/ as focused vitest suites, including vectors.test.ts (encode

  • decode every shared conformance vector), istream.chunked.test.ts (every vector fed one byte at a time), skip-ids.test.ts (every vector that carries skip_ids, decoded by a receiver that ignores those ids at every nesting level — contiguous, one byte at a time, and split in two at every byte boundary), heap-free-codec.test.ts (no allocation primitive on a codec path, a flat heap over encode and decode, no view into a fed or one-shot buffer) and pooled-decoder-state.test.ts (a decode aborted at every cut point leaves nothing behind for the next one).

The vector-driven suites each print one summary line — [vectors] 131 vectors, none gated out by requires, 524 checks — so a run says how much of the shared suite it actually executed, and a file that arrived truncated or a group gated out by requires shows up as a smaller number rather than as silence. This port compiles no feature out, so nothing is ever gated.

assets/test_vectors.json carries six blocks and this port runs all six: vectors, invalid_utf8, sequence_growth, header_limits, header_limits_nested and boolean_tolerant. The file is a verbatim copy of the one in corelib-c-cpp, which authors it. A daily CI job (.github/workflows/shared-vectors.yml) compares this copy's sha256 against that file on corelib-c-cpp@main, so a copy left behind by an upstream change is reported rather than going unnoticed.

sequence_growth holds the wrapper-array growth cases of §7.2 item 8, replayed by sequence-growth.test.ts for both element kinds at three chunkings. This port declares dynamic_arrays: its wrapper-array containers are JS arrays that grow at decode time, so the block applies. The cases are cap-relative and the run installs max_dyn_array_count = 8. Growth geometry splits in two: the backing store's reallocation strategy is the engine's amortised doubling, which is not this port's to pin, while the fill is, and is asserted as one write per slot in a single pass.

header_limits holds the truncated over-ceiling headers of §6.2.1 / §6.3 — bytes that declare a length or an element count and then end, with no payload behind them — replayed by header-limits.test.ts. The ceiling answers at that word, before the payload is asked for, so the verdict is terminal and never INCOMPLETE; which ceiling the case configures decides the category, INVALID for a schema maxlen and LIMIT_EXCEEDED for a §6.2.1 receiver cap. This port declares receiver_caps: its generated layer carries receiver caps distinct from schema bounds, compared inside the visitor's own fixlenBegin / arrayBegin, which this library raises at the header word. Every rejection case is paired with an in-cap control that must still answer incomplete, and the block is also run with the ceilings lifted — where all six rejections fall back to incomplete, since this port's FIXLEN_MAX (INT32_MAX) sits above even the amplification case's 1 GiB claim.

header_limits_nested is that same assertion one or two sequence frames deeper, replayed by header-limits-nested.test.ts — the axis the flat block leaves untested, since every case there sits at the top level. Its cases open a sequence, declare the over-ceiling word inside it and end with the frame still open, so a decoder has a second, independent reason to answer incomplete and a port that binds its ceiling to the top-level scope looks plausible while capping nothing. The runner descends the frames chain outermost first, applies the same leaf the flat block uses (test/helpers/header-limits.ts) at the innermost depth, and runs a negative control over every rejection with that case's own kind of ceiling lifted to 65536: all four then answer incomplete instead, which is what shows the verdict was the ceiling's rather than an unclosed frame's. All 8 cases run, none gated.

boolean_tolerant holds §4.4's decode half — bytes carrying 2, 256 or 2^64-1 at a boolean position — replayed by boolean-tolerant.test.ts. No conforming encoder emits such a value, so the positive vectors cannot reach the rule at all: those bytes only ever arrive from someone else's encoder. A boolean carries no width bound, so every non-zero value is true — not INVALID, and not a truncation to false — and the decode is normalized, which only the re-encode makes visible. Each case is therefore asserted three ways: the outcome is complete, the destination holds exactly 0 / 1 (poisoned with 0xaa first, so a decoder that never writes cannot pass), and the re-encode of what was decoded is byte-compared against the block's reencoded_hex — 0001, never 0002. An array decodes through the bool destination, which is where this library performs the normalization; a scalar boolean has no callback of its own and arrives on unsigned with its full 64 bits, so the !== 0 test there is the generated layer's and the test performs it, after asserting that the delivered value and its two halves agree. In this block an unsatisfied requires tag means the message must be rejected, not skipped (§4.4 lifts the width bound the type carries, never the one a narrowed build has) — a path this port never takes, since it compiles no feature out.

CI type-checks, tests and builds on Node 20 / 22 / 24 / 26, smoke-tests the bundle on Node, Deno and Bun, and publishes coverage badges; a separate docs.yml deploys the TypeDoc API reference to GitHub Pages.

Benchmarks

Three standalone tools, specified by BENCH_SPEC.md and mirrored in every other port — same datasets, same timing rules, same output grammar — so the numbers compare directly across languages:

npm run perf              # per-op cost on the 170-byte perf message
npm run bench             # throughput table (MB/s) over the four shared datasets
npm run bench:callgrind   # machine-independent instructions/op under Valgrind

bench prints ten rows over four datasets: a 1000-element u64 array, the small typical message, an unbounded 1 MB blob, and a composite message that reaches what the flat ones miss — a 64-element wrapper array, 320 bytes of 1-, 2-, 3- and 4-byte UTF-8, nesting three deep, a default-valued field the encoder must not write, and a two-byte field header. Three of the encoded sizes are cross-port parity checks (perf = 170 bytes, blob 1MB = 1,000,005, composite = 956); test/bench-datasets.test.ts holds the datasets to them, and test/bench-grammar.test.ts holds the tools to the output grammar.

Every encode row writes into a caller-supplied buffer rather than the accumulator. The blob 1MB rows are the ones that exercise streaming end to end: one-shot is a single contiguous write into a 1,000,005-byte buffer, streaming is the same bytes through a 4096-byte buffer with a flush sink (~245 flushes), and decode: blob 1MB is fed back in 4096-byte chunks. The difference between the two encode rows is what the divisible-run flush path costs, and it is legible only under Ir/op. BENCH_SPEC's optional blob 1MB passthrough row is absent: pass-through is forbidden (§5.1.6), so every string / blob run is copied through the output buffer. Both copies are TypedArray.set — the whole payload when it fits, a range per flush when it does not, the latter through one of the itemised §6.6.2 handles under Memory handling.

Since JS engines expose no portable cycle counter, perf uses CPU time/op as the code-cost proxy; bench:callgrind counts instructions/op under Valgrind (two rep counts per workload, subtracted, on a --predictable V8) for a fully machine-independent figure. The same tools under Node (V8) and Bun (JavaScriptCore) give directly comparable numbers. tsx bench/bench.ts --smoke runs every row exactly once — a liveness check for the rows, never a measurement.