base85n
v0.5.1
Published
Base85N: a compact encoding for data embedded in text formats (JSON, XML, HTML, config), with Dynamic Passthrough and Fill modes
Maintainers
Readme
base85n
A TypeScript implementation of Base85N, an encoding for data that has to be
embedded in a text-based format — JSON, XML, HTML attributes, configuration files
— where Base64 would otherwise be used and the size or the cleanliness of the
result matters. It uses a single 85-character alphabet (Alphabet-N) that excludes
", ', \, <, >, & and all whitespace, so output goes into a JSON
string, an XML node or a quoted attribute with no second escaping layer, with an
adaptive Dynamic Passthrough (DP) mode for efficient, partially human-readable
representation of compatible byte sequences. See
the specification
for the full normative text, in particular Section 4.2's eight replacement
alphabets and Section 6.1's single-scan Dynamic Passthrough
encoding procedure, which this package follows exactly.
Before you embed the output: four containers need the value quoted. Alphabet-N does contain
`,$,{and=— free in JSON, XML and HTML, not free everywhere. So encoded output must not be pasted raw into a JavaScript template literal (`ends it,${interpolates — both occur in ordinary output, a backtick about one character in 85), an unquoted HTML attribute, a plain YAML scalar, or a double-quoted shell word. An ordinary'…'or"…"string is always safe. Full table, checked against real parsers: Embedding: where the output can be pasted verbatim.
Install
npm install base85nOr, to work on this package inside a clone of the repository, npm install with
no arguments from this directory.
Build
npm run buildCompiles src/ to dist/ (ESM, with .d.ts type declarations) using tsc.
Test
npm testRuns the vitest suite under test/, which:
- Verifies every golden vector in
../testvectors/vectors.jsonround-trips through bothencodeanddecode. - Runs seeded (deterministic) random round-trip property tests across a wide range of input lengths and byte compositions (raw random bytes, Alphabet-N literals, R-Set characters, and donor characters).
- Exercises explicit edge cases: empty input, 1-4 byte inputs, the
MIN_PASSTHROUGH_BYTESboundary, multi-segment DP output (> MAX_DP_OUTPUT_CHARS_PER_SIGNALcharacters), theMAX_DP_ANALYSIS_BYTESwindow boundary, and every byte value 0-255. - Asserts
decode()throws aBase85NDecodeError(never returns garbage) on a variety of malformed inputs.
Usage
import { encode, decode, Base85NDecodeError } from "base85n";
const bytes = new Uint8Array([72, 101, 108, 108, 111]);
const text = encode(bytes); // => Base85N-encoded string
try {
const roundTripped = decode(text); // => Uint8Array, equal to `bytes`
} catch (err) {
if (err instanceof Base85NDecodeError) {
console.error(err.code, err.message);
}
}API
function encode(data: Uint8Array): string;
function decode(s: string): Uint8Array; // throws Base85NDecodeError
class Base85NDecodeError extends Error {
readonly code:
| "invalid_character"
| "unexpected_end_of_stream"
| "undefined_signal"
| "invalid_final_block";
readonly position: number | undefined; // index into the whitespace-stripped input, if known
}Also exported: the ALPHABET_N_CHARS_STR, MIN_PASSTHROUGH_BYTES,
MAX_DP_OUTPUT_CHARS_PER_SIGNAL, MAX_DP_ANALYSIS_BYTES and
REPLACEMENT_ALPHABETS from
the specification.
Versioning
The major and minor version track the specification version this package
implements — 0.5.x implements specification v0.5.0, whose wire format is
frozen. The patch level is this package's own: packaging and documentation
fixes that change no encoded output. Anything that would change the wire format
would change the specification's version first.
