@pixagram/turbobase64
v1.0.1
Published
Fast, strict base64 for JavaScript: the engine's native Uint8Array.toBase64/fromBase64 where they exist, a scalar asm.js kernel everywhere else, plus a SIMD128 (SIMD.js-style) transcription of the Muła–Lemire vectorized algorithm.
Maintainers
Readme
@pixagram/turbobase64
Fast, strict base64 for JavaScript, with no dependencies and no build step.
- Native first. When the engine has
Uint8Array.prototype.toBase64/Uint8Array.fromBase64(current browsers, Node 26+), the module uses them — after probing that they behave as expected — and gets several GB/s. - asm.js fallback. Everywhere else (Node 22 and 24 LTS, older browsers) a scalar asm.js kernel takes over: 0.8–1.7 GB/s where asm.js is compiled to WebAssembly, 0.3–0.9 GB/s as plain JavaScript.
- Strict decoding. Anything outside
A-Za-z0-9+/and well-formed=padding throws aTurboBase64Errorwhose.indexpoints at the offending character. Padding is optional. - A SIMD128 kernel for the curious. The Muła & Lemire vectorized algorithm (arXiv:1704.00605), transcribed instruction for instruction onto a minimal SIMD.js-style layer that lives inside the same asm.js module. It is slower than the scalar kernel — see Performance for why — and kept for study and benchmarking.
Install
npm install @pixagram/turbobase64Usage
const TurboBase64 = require('@pixagram/turbobase64'); // CommonJS
import TurboBase64 from '@pixagram/turbobase64'; // ESM (default import)
// <script src="turbobase64.js"></script> defines window.TurboBase64
const b64 = new TurboBase64(); // picks 'native' or 'scalar' for this engine
b64.kernel; // 'native' | 'scalar' | 'simd'
const bytes = new TextEncoder().encode('Fast and efficient Base64 encoding!');
const text = b64.encode(bytes); // 'RmFzdCBhbmQgZWZmaWNpZW50IEJhc2U2NCBlbmNvZGluZyE='
const back = b64.decode(text); // Uint8Array, equal to bytes
b64.decode('RmFzdA'); // unpadded input is fine
try {
b64.decode('Rm*zdA==');
} catch (e) {
e instanceof TurboBase64.TurboBase64Error; // true
e.index; // 2
}
TurboBase64.encode(bytes); // one-off helpers on a shared instance
TurboBase64.decode(text);Kernels
new TurboBase64({ kernel })
| kernel | what runs |
|---|---|
| auto (default) | the engine's toBase64 / fromBase64 if present and sane, otherwise scalar |
| native | the engine implementation; the constructor throws if it is not available |
| scalar | the scalar asm.js kernel: one 12-bit → two-character table for encoding, four pre-shifted 256-entry tables for decoding, a sign bit as the error flag |
| simd | the SIMD128 transcription of the paper's SSE code |
TurboBase64.nativeAvailable tells you what auto will do. The native probe is
a semantics check, not a typeof: it encodes a known vector, decodes it back,
checks that unpadded input is accepted and that invalid input throws, so a
partial polyfill is never trusted.
A native instance does not build the asm.js module or its 2 MiB heap unless
you touch .asm / .heap / .U8, which exist for the SIMD128 operations and
the tests.
API
new TurboBase64(options?)—options.kernelas above.encode(bytes) → string—bytesis aUint8Array, anyArrayBufferViewor anArrayBuffer. Standard alphabet,=padding.encodeToBytes(bytes) → Uint8Array— the same output as ASCII bytes.decode(base64) → Uint8Array— throwsTypeErrorfor non-strings andTurboBase64Errorfor invalid input.decodeBytes(base64Bytes) → Uint8Array— the input already as ASCII bytes.kernel,requestedKernel— resolved and requested kernel names.TurboBase64.encode(bytes),TurboBase64.decode(base64)— a lazily created shared instance.TurboBase64.nativeAvailable,TurboBase64.KERNELS,TurboBase64.ALPHABET,TurboBase64.TurboBase64Error.TurboBase64.TurboBase64Module,TurboBase64.LAYOUT— the raw asm.js module and its heap map, for reuse of the SIMD128 layer.
TypeScript declarations ship in turbobase64.d.ts.
What decode accepts
| input | result |
|---|---|
| canonical, padded | decoded |
| padding omitted (Zm9v, Zg) | decoded |
| = in the wrong place, too many, or a length ≡ 1 mod 4 | TurboBase64Error, .index at the padding |
| any character outside the alphabet, including - _ (URL-safe) and non-ASCII | TurboBase64Error, .index at the character |
| ASCII whitespace inside the input | scalar / simd: rejected. native: skipped — Uint8Array.fromBase64 ignores whitespace by specification, and pre-scanning every input to reject it would cost more than the native call saves. If you must reject whitespace, use kernel: 'scalar'. |
Errors from the native path carry the same message and .index as the asm.js
kernels; the position is located only when an error actually happens.
Performance
Median throughput at 1 MiB of random data, MB/s of binary data (encode / decode), same x64 sandbox. Chrome runs the asm.js module as plain JavaScript; both Node versions compile it to WebAssembly; Node 26 also has the native functions, Node 22 does not.
| implementation | Chrome 153 | Node 26 | Node 22 |
|---|---:|---:|---:|
| auto (→ native where present) | 2 130 / 932 | 13 300 / 7 800 | = scalar |
| scalar asm.js kernel | 403 / 279 | 1 211 / 1 196 | 730 / 903 |
| simd128 asm.js kernel | 53 / 49 | 150 / 114 | 124 / 116 |
| previous @pixagram/turbobase64 0.0.2 class | 127 / 26 | 134 / 35 | 123 / 28 |
| base64-js | 93 / 256 | — | — |
| Node Buffer | — | 12 672 / 14 019 | 11 642 / 13 684 |
The auto wrapper adds nothing measurable over calling toBase64 /
fromBase64 yourself. Numbers move a lot with the machine, the tab and garbage
collection; read the pattern, not a cell.
Three things worth knowing:
- asm.js ahead-of-time compilation is leaving V8. V8 deprecated asm.js validation in 2026 (Chrome 153 no longer has it; Node 26's V8 14.6 still does). Valid asm.js remains good, allocation-free, monomorphic JavaScript, so the scalar kernel still beats the other pure-JS encoders here — but it is the native path that guarantees native speed.
- Why the SIMD128 kernel is slow. The paper's algorithm is built around
pshufb, the one instruction with no 32-bit shortcut: 16 dependent byte lookups per call, two calls per 12 encoded bytes, four per 16 decoded characters. Emulated, that alone costs more than a whole scalar block. The rest of the SIMD128 layer is 4 or 8 word operations and is cheap. In JavaScript the paper's ideas pay off as scalar SWAR and pre-shifted tables, not as emulated vectors; real SIMD base64 in the browser is WebAssembly SIMD (i8x16.swizzleis pshufb). - Decode is stricter than before. The 0.0.2 class silently produced bytes for invalid input and returned wrong bytes for the URL-safe alphabet; this version throws in both cases. Whitespace handling differs per kernel as described above.
How it works
One "use asm" module holds both layers, because asm.js code can only reach
another module through the slow foreign FFI and a 12-byte block needs ~14
vector operations.
SIMD128 — registers are heap offsets. A v128 is a 16-byte-aligned offset
into the module heap; every operation is op(dst, src, src). Ops with a cheap
32-bit formulation run SWAR-style on the Int32 view, four lanes at a time:
| op | SSE | how |
|---|---|---|
| v128_and/or, v128_testz | pand / por / ptest | 4 × Int32 |
| u32x4_shr | psrld | 4 × >>>, bits cross bytes exactly like hardware |
| i8x16_add | paddb | ((a&L)+(b&L)) ^ ((a^b)&H) |
| i8x16_eq | pcmpeqb | zero-byte detection, spread with imul(m>>>7, 255) |
| i8x16_gt | pcmpgtb | signed compare via sign flip + unsigned >= (below) |
| u8x16_subs | psubusb | byte subtract masked by unsigned >= (below) |
| i8x16_swizzle | pshufb | 16 byte lookups, high bit of the index zeroes the lane |
| u16x8_mulhi, i16x8_mul | pmulhuw / pmullw | 8 × imul |
| i16x8_maddubs, i32x4_madd | pmaddubsw / pmaddwd | pairs of imul, saturation for the former |
With H = 0x80808080 and L = 0x7f7f7f7f, t = (a|H) - (b&L) is
128 + (a&127) - (b&127) per byte and never borrows across bytes, so bit 7 of
each byte says (a&127) >= (b&127); then ge = ((a & ~b) | (~(a^b) & t)) & H
marks the bytes where a >= b unsigned, psubusb = (t ^ (~(a^b) & H)) & spread(ge)
and pcmpgtb(a, b) = spread(~ge(b^H, a^H) & H).
Kernels. encodeSimd / decodeSimd follow the paper's SSE code line by
line: encode is pshufb [b1 b0 b2 b1] → pand / pmulhuw for values a, c →
pand / pmullw for b, d → por → the psubusb 51 / pcmpgtb 26 / pshufb
offsets / paddb translation (Table IV); decode is psrld 4, pand 0x2F, the
lut_lo / lut_hi / lut_roll lookups with ptest validation (Algorithm 3), then
pmaddubsw 0x01400140 → pmaddwd 0x00011000 → pshufb pack (Fig. 7). The
scalar kernels share the same heap and I/O path, so the two can be compared
like for like.
Heap (2 MiB). Constants at 0x000, an eight-register file at 0x100, the
scalar tables at 0x300–0x3fff, a 768 KiB input window at 0x4000 and a
1 MiB output window at 0xC4100. Larger inputs stream through the windows in
multiples of 12 (encode) or 16 (decode) bytes, so only the last window has a
tail. Strings cross into the heap with TextEncoder.encodeInto and out with
TextDecoder.
Development
npm test # node turbobase64.test.js — SIMD128 ops vs per-lane SSE semantics, both kernels
# vs Buffer at every length 0–600, window boundaries, padding rules, error
# indices; the native kernel is tested too when the engine has it
npm run bench # node --allow-natives-syntax bench.js — kernels vs the 0.0.2 class vs Buffer
npm run build:html # regenerates test-suite.html from test-suite.template.htmltest-suite.html is a self-contained page for browsers: a correctness matrix
(this module's auto, simd and scalar kernels, the 0.0.2 class, base64-js,
Uint8Array.toBase64 and btoa/atob, all against an in-page oracle), the
SIMD128 operation tests, and a benchmark with adaptive batching so the ~100 µs
clamp on performance.now() cannot skew small sizes. Open it from disk.
reference/ keeps the 0.0.2 class and base64-js verbatim for those comparisons;
neither is part of the npm package.
Bundling
There is nothing to transpile: the module is plain ES2020 CommonJS with a
UMD-style footer. If you minify it, make sure the minifier leaves the
"use asm" function alone (Terser and UglifyJS do by default); rewriting its
body breaks asm.js validation on engines that still perform it, and the code
then runs as ordinary JavaScript.
License
MIT. This implementation was written from the paper's description; it contains no code from the TurboBase64 C library.
