cemalloc
v0.1.1
Published
A userspace slab/page allocator and reusable binary-memory pool for JavaScript and Node.js
Downloads
765
Maintainers
Readme
cemalloc
A userspace slab/page allocator and reusable binary-memory pool for JavaScript and Node.js.
cemalloc is an advanced application-level memory manager inspired by jemalloc's slab and arena architecture, designed for Node.js workloads that allocate large numbers of short-lived Uint8Array / ArrayBuffer instances (such as database storage engines, B+Tree page managers, binary parsers, serialization libraries, and networking protocols).
What is this?
cemalloc is an explicit memory management runtime for Node.js. It does not replace V8's internal allocator or fake C's malloc. Instead, it pre-allocates large, contiguous ArrayBuffer memory regions (arenas) and returns reusable Uint8Array views over them via size-segregated slabs, dedicated page pools, and scratch bump allocators.
The Problem It Solves
In Node.js, repeatedly calling new Uint8Array(size) in high-throughput workloads creates hundreds of thousands of short-lived ArrayBuffer objects in V8's C++ heap:
- Garbage Collection Pressure ("Stop-The-World"): V8's GC must continuously pause JavaScript execution to clean up short-lived buffer allocations.
- Tail Latency Spikes (p99): While average allocation time appears low, the 99th percentile latency degrades severely due to Major GC cycles and external memory adjustments.
- Memory Bloat & Fragmentation: Constant buffer churn leaves V8 external memory fluctuating wildly, preventing predictable memory consumption.
How cemalloc Solves It
- Zero-Allocation Recycling:
cemallocreserves contiguous memory upfront (16 MB default chunks) and slices them into uniform 64 KB slabs. - Instant $O(1)$ Slab & Page Reuse: Both
alloc()andfree()operate in constant time using direct lookup tables (Int32Array) and intrusive LIFO free-lists. - Bypassing the V8 GC: Calling
allocator.free(buffer)immediately returns the slot to an internal free-list. The V8 Garbage Collector is never invoked, cutting GC pauses by up to 84% and boosting throughput by 3x to 12x for buffers ≥ 256 bytes.
Key Features
- Multi-Chunk Arenas: Pre-reserves 16 MB
ArrayBufferregions and allocates additional backing chunks on-demand without invalidating active views. - Uniform 64 KB Slabs: Slices chunks into 64 KB slabs with bitwise $O(1)$ address resolution (
buffer.byteOffset >>> 16). - Direct Lookup Table (LUT): Employs an
Int32Array(16385)LUT to resolve size classes in a single CPU instruction without linear scans. - Scratch & Bump Allocators: High-throughput bump-pointer arena (
allocator.scratch()) with $O(1)$ bulk deallocation (reset()) and checkpoints (mark()/rewind()). - Slab Reclamation & Trimming: Tracks
EMPTY,PARTIAL, andFULLslabs;allocator.trim()reclaims unused slabs and safely dereferences empty chunks. - Memory Pressure Policies: Configurable
softLimit(triggers trimming and callbacks) andhardLimit(prevents uncontrolled memory growth). - 32-Bit Packed Handles & Generational Protection: Low-level integer handles encoding Chunk, Slab, Slot, and Generation counters to detect stale handles and double-frees.
- Database Page Cache: Reusable page caching layer on top of
allocPage()with pin counts, dirty flags, and LRU eviction. - Adaptive Size Classes: Analyzes allocation size frequency distributions to suggest intermediate classes, reducing internal waste by up to 38%.
- Allocation Lifetime Profiling & Hot/Cold Slabs: Assigns monotonic operation epochs to analyze lifetime distributions (
<10,10-100,1k,>10k), with semantic lifetime hints (allocator.temp(),allocator.persistent()). - SharedArrayBuffer & Worker Threads: Supports
backend: 'shared', worker-local arenas, and lock-freeRemoteFreeQueuecross-worker deallocation. - Invariant Checker: Development validator (
allocator.verify()) that verifies freelist integrity, slot counts, and state invariants.
Installation & Quick Start
git clone https://github.com/litepacks/cemalloc.git
cd cemalloc
npm install1. Standard Slab Allocation
import { Allocator } from "./src/index.js";
const allocator = new Allocator({
initialSize: 16 * 1024 * 1024,
classes: [64, 128, 256, 512, 1024, 4096, 16384]
});
// Allocates 180 bytes -> automatically selects 256-byte class
const buffer = allocator.alloc(180);
buffer[0] = 42;
// Free buffer back to pool
allocator.free(buffer);2. Scratch Bump Allocator (Fastest for Parsers & Queries)
const scratch = allocator.scratch(4 * 1024 * 1024);
const a = scratch.alloc(128);
const mark = scratch.mark(); // Checkpoint
const temp1 = scratch.alloc(1024);
const temp2 = scratch.alloc(2048);
// Rewind back to checkpoint
scratch.rewind(mark);
// Bulk reset in O(1)
scratch.reset();3. Dedicated Page Allocator (Database B+Trees)
const page = allocator.allocPage();
page[0] = 0xAA;
allocator.freePage(page);4. 32-Bit Packed Handles & Generational Protection
const handle = allocator.allocHandle(256);
const view = allocator.getView(handle);
allocator.freeHandle(handle);
// Attempting to use a freed handle throws StaleHandleError
// allocator.getView(handle); // Throws StaleHandleError!5. Slab Trimming & Reclamation
// Reclaim completely empty slabs and dereference empty non-primary chunks
const report = allocator.trim();
console.log(report);
// { slabsReleased: 2, chunksReleased: 1, bytesDereferenced: 16908288 }6. Invariant Verification
// Validates all internal invariants (used + free == total, freelist integrity, bounds)
allocator.verify(); // Returns trueTest & Benchmark Commands
# Run full Vitest suite (18 test files, 56 tests)
npm test
# Run unified high-precision benchmark suite
npm run bench
# Run pool-vs-slab decomposition experiment
npm run bench:decomp
# Run real B+Tree storage engine benchmark
npm run bench:storage
# Run multi-wave memory plateau test
npm run bench:plateau
# Run continuous soak test
npm run bench:soak -- --duration=10s
# Run multi-worker thread scaling benchmark
npm run bench:workers
# Run adaptive size class analysis benchmark
npm run bench:adaptiveIn-Depth Experimental Findings
A complete empirical study answering the 13 core research questions is published in BENCHMARK_ANALYSIS.md.
Summary of Results:
- Microbenchmarks (256 B – 16 KB):
cemallocis 3.1x to 11.9x faster than nativenew Uint8Array(n), reducing p99 tail latency by up to 84%. - Scratch Bump Allocator: Reaches 18.9M ops/sec with $O(1)$ bulk reset, outperforming general slab allocation for query and parser scopes.
- Storage Engine B+Trees: Delivers 5.9x throughput over native allocation and completely eliminates 25+ MB of ArrayBuffer GC churn.
- Data-Oriented Metadata: TypedArray metadata runs 3.1x faster and generates 18x less heap garbage than JS object-based metadata.
- Where Native Wins: For small buffers (≤ 64 bytes), native V8 allocation is ~26% faster due to young-generation nursery inlining.
License
MIT
