npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@mizchi/jsimd

v0.3.0

Published

Small WebAssembly SIMD kernels, resident data structures, and dynamic Float32 fusion for JavaScript

Readme

@mizchi/jsimd

Small WebAssembly SIMD kernels, Wasm-resident data structures, and an opt-in dynamic Float32 fusion compiler for data-parallel JavaScript hot paths.

Goal

jsimd provides small, tree-shakeable data structures and bulk operations powered by hand-written WebAssembly SIMD kernels. It targets JavaScript hot paths where explicit SIMD can outperform the equivalent built-in implementation end to end.

Each feature ships as an independent package subpath with documented performance, boundary-cost, and bundle-size trade-offs.

Install

pnpm add @mizchi/jsimd

Runtime requirements

  • Node.js 24.5 or later
  • Deno 2.6 or later
  • Vite 8 or later
  • WebAssembly SIMD
  • WebAssembly ESM Integration for direct .wasm imports

The optional shared-buffer entrypoint additionally requires WebAssembly threads and shared memory. The ultra-log-log-parallel entrypoint requires SharedArrayBuffer and Workers but performs no atomic Wasm access. Browsers normally expose shared buffers through cross-origin isolation. Node uses node:worker_threads; Deno and browsers use Web Workers. Their factories are asynchronous because each Worker must initialize its module.

All supported environments use the same entrypoints. Existing single-threaded entrypoints expose synchronous operations after ESM Integration instantiates their Wasm modules. shared-buffer is the exception: its asynchronous factories compile the same kernel in each Worker against caller-owned shared memory. Consumers must also enable explicit resource management (using / Symbol.dispose).

f32-fusion is another explicit exception: it ships no static Wasm asset and asynchronously emits and compiles a restricted SIMD expression at runtime. It requires CSP permission for WebAssembly.compile(); use supportsF32Fusion() when deployment policy is unknown.

Usage

import { indexOf } from "@mizchi/jsimd/bytes";
import { WaveletMatrixUint32 } from "@mizchi/jsimd/wavelet-matrix-uint32";

const bytes = new TextEncoder().encode("simd:simd");
indexOf(bytes, 0x3a); // 4
indexOf(bytes, new TextEncoder().encode("simd")); // 0

// Wasm-resident data structures own memory, so release them with `using`.
using values = WaveletMatrixUint32.from([3, 1, 4, 1, 5, 9]);
values.rank(1, values.length); // 2
values.rangeFreq(0, values.length, 1, 5); // 4
values.quantile(0, values.length, 3); // 4

The single-threaded entrypoints use direct Wasm ES module imports supported by current Vite and Deno. Module loading performs initialization, after which their exported operations are synchronous.

Implementation guide

Each link below opens the feature README with its complete API, algorithm and literature sources, benchmark setup, and exact results. This table is only a rounded selection guide.

Owning structures keep data in Wasm linear memory and should be declared with using. They are a good fit when several bulk operations reuse resident data. Native arrays and collections usually remain the better choice for small inputs, one-shot work, point access, or workloads that repeatedly materialize complete results back into JavaScript.

Disposal returns allocator blocks to an internal reuse pool; it does not shrink WebAssembly.Memory. A stable reservedBytes or memoryBytes value after a workload is therefore expected. Stateful entrypoints expose allocatorStats() so tests can require liveAllocations and liveBytes to return to their baseline after each using scope.

Data structures

The canonical bitmap subpaths are bitmap, rank-select-bit-vector, and roaring-bitmap. bit-vector is intentionally reserved for a future immutable packed-bit sequence without a rank/select index; it is not currently exported. Pre-announcement compatibility aliases were removed so each structure has one public name.

| export | purpose | observed speedup | trade-off | minified JS + Wasm, gzip | | :--------------------------------------------------------------------- | :------------------------------------------- | :--------------- | :---------------------------------------------------- | :----------------------- | | adaptive-simd-page-i32 | Adaptive frozen i32 pages/columns | 0.5–41x | Compressed decode/construction can be slower | 5.84 kB + 1.63 kB | | bitmap | Growable and fixed dense mutable bitmaps | 9.8–19.8x | Small and point-heavy cases were not measured | 2.94 kB + 0.26 kB | | bit-histogram32 | Streaming positional popcount for u32 flags | 1.9–14.4x | One word was 9.03x slower than JavaScript | 2.17 kB + 0.27 kB | | bit-matrix | Dense Boolean matrix and frozen CSR | 6.56x dense | CSR traversal and small matrices can be slower | 3.29 kB + 0.41 kB | | byte-key-flat-hash | Variable-byte-key map with resident arena | 2.00x bulk | Individual gets were 12.5x slower | 3.37 kB + 0.73 kB | | compressed-string-table | Frozen front-coded byte strings | 2.00x byte eq | Decode and random materialization can be slower | 4.18 kB + 0.34 kB | | columnar | Shared-mask i32/u32/u8 column predicates | 2.9–40.5x | Dense extraction and one-shot build can lose to JS | 7.17 kB + 1.16 kB | | binary-vector-index | Hamming search and exact PDX rerank | 6.5–9.8x | Recall depends on candidate count | 4.17 kB + 0.43 kB | | blocked-vector-array | Repeated exact Float32 vector scoring | 1.2–9.2x | Construction was 7.24x slower than typed-array copy | 2.87 kB + 0.92 kB | | bit-sliced-column | Repeated predicates over static u8 columns | 17.6–29.6x | Construction, point reads, and small scans excluded | 3.00 kB + 0.43 kB | | blocked-bloom-filter | Bulk negative filter before exact lookup | 1.6–5.5x E2E | All-hit lookup was 1.22x slower | 2.49 kB + 0.37 kB | | elias-fano-sequence | Global/partitioned monotone sequences | 0.03–11.2x | Point access, decode, and uniform partitions are slow | 5.09 kB + 0.86 kB | | f32-vector | Resident Float32 vector operations | 4–24x bulk | Tiny and one-shot work can be slower | 2.12 kB + 0.31 kB | | flat-hash | Batched u32/u64 hash set/map | 6.8–12.0x | Individual has and get calls are slower | 3.35 kB + 0.95 kB | | flat-hash-fixed16 | UUID/hash-keyed map and set | 25.6–27.9x bulk | Numeric and point-key workloads can be slower | 3.02 kB + 0.62 kB | | fingerprint-group16 | SwissTable control groups and tables | 1.6–4.9x bulk | Individual probes are slower | 2.33 kB + 0.27 kB | | fm-index-bytes | Frozen full-text byte search | 6.86x count | Construction, locate, and mutable text can be slower | 4.99 kB + 1.46 kB | | i32-array | Resident fixed i32 arrays | 3–7x | One-shot and sub-1K cases were not measured | 2.21 kB + 0.34 kB | | matrix2d | Resident Float32 matrix multiplication | ~1–9.5x | 4×4 was near parity; BLAS/GPU were not compared | 2.51 kB + 0.23 kB | | matrix3d | Resident batched matrix multiplication | 5.0–7.3x | Compared only with resident generic JS loops | 2.67 kB + 0.27 kB | | packed-delta-uint32-list | Compressed postings and monotone lists | 0.06–1.4x | Full decode and lower-bound queries are slower | 2.91 kB + 0.88 kB | | rank-select-bit-vector | Frozen indexed RankSelectBitVector | 1.5–3.0x bulk | Single-query rank is slower | 2.97 kB + 0.77 kB | | roaring-bitmap | Compressed mutable u32 bitmap | 2.2–175x | Construction and point-heavy cases were not measured | 4.74 kB + 0.53 kB | | shared-buffer | Shared memory, queues, snapshots, reduction | 1.30–7.96x bulk | Pool lease was 1.10x slower; scheduling excluded | 26.65 kB + 0.33 kB | | static-mphf-u32 | Frozen perfect hash for known u32 keys | 1.75x bulk | Individual lookup and construction are slower | 4.09 kB + 0.39 kB | | ultra-log-log | Mergeable approximate u32 distinct count | 2.16–22.53x | Small batches select scalar JavaScript | 4.08 kB + 0.57 kB | | ultra-log-log-parallel | Persistent-Worker bulk distinct count | 2.39–5.38x E2E | Forced 4K Worker path was 3.45x slower | 9.43 kB + 0.57 kB | | wavelet-matrix-uint8 | Rank/range queries over frozen bytes | 4.8x vs u32 | Direct byte access is slower | 4.30 kB + 1.11 kB | | wavelet-matrix-uint16 | Range queries over frozen u16 sequences | 2.3–329x | Direct access and exact rank are slower | 4.30 kB + 1.11 kB | | wavelet-matrix-uint32 | Range statistics over frozen u32 sequences | 2.2–100x+ | Direct access and exact rank are slower | 4.24 kB + 1.09 kB |

Stateless and copy-inclusive kernels

| export | purpose | observed speedup | trade-off | minified JS + Wasm, gzip | | :--------------------------------- | :---------------------- | :--------------- | :---------------------------- | :----------------------- | | bytes | General byte operations | 4.9–18.1x | Small inputs use JS fallbacks | 1.40 kB + 0.77 kB | | endian | Batched u32 decoding | 1.0–2.2x | Small inputs were at parity | 1.11 kB + 0.18 kB |

Dynamic kernels

| export | purpose | observed speedup | trade-off | minified JS + Wasm, gzip | | :----------------------------------------- | :----------------------------------- | :--------------- | :---------------------------------------------- | :----------------------- | | f32-fusion | Runtime-generated fused Float32 pass | 3.76x resident | Cold compile, resident memory, and CSP required | 3.34 kB + 0 kB |

How to read the numbers

Performance results were recorded on Apple M5 with Node 24 or Deno 2.6 as documented by each linked feature README. They are workload samples, not cross-feature scores. Resident benchmarks usually exclude construction and final materialization; copy-inclusive kernels explicitly include boundary copies. Rerun the linked benchmark on the target engine and data distribution before choosing a representation.

Admission is based on the documented primary workload, not on every convenience method. A row that reports a slower JS case is retained only when a separate, representative bulk workload wins; that slower operation is outside the performance contract.

Build sizes come from isolated Vite 8.2 production fixtures. Each cell reports the gzip sizes of the minified JavaScript output and its independently emitted, stripped Wasm asset. f32-fusion reports zero Wasm because it generates the module at runtime. Importing a subpath does not pull in another feature's Wasm. A real application may share wrapper code or add other runtime code, so these figures are marginal fixtures rather than a prediction of total bundle size. Raw sizes remain available in each feature README.

The package is distributed as one npm package with subpath exports. npm releases contain compiled JavaScript and adjacent declarations; consumers do not need TypeScript runtime transformation for files under node_modules. Each .wasm import is typed by an adjacent kernels.d.wasm.ts, following Vite's allowArbitraryExtensions convention, without a generated environment-specific loader.

Examples

  • multithread-vector-search shards an exact Float32 index across Web Workers, uses shared SPSC task notification and result slots, and merges only workerCount × k candidates on the coordinator.

Storage package

@mizchi/jsimd-columnar builds typed schemas, ZoneMap/projection pushdown, resident SIMD page caching, and interchangeable Memory, IndexedDB, and Node filesystem backends. It is versioned separately as an experimental 0.x package: warm queries win the recorded selective workload, while cold storage restore remains slower than already-resident page-aware JavaScript.

Development

Implementation priorities, benchmark gates, and deferred candidates are maintained in TODO.md. A kernel is retained only when a documented workload justifies its Wasm boundary and bundle cost against the best relevant JavaScript builtin or scalar reference.

pnpm install
just build
just test
just bench
just memory-profile

just build compiles 30 Zig 0.16 src/<name>/kernels.zig sources and the deliberately hand-written bytes/kernels.wat source with SIMD128 and wasm-opt -Oz. Generated Wasm files are checked for their public exports, expected SIMD instructions, and per-module size limits, then validated with wasm-tools. just build-package emits the npm payload into packages/jsimd/dist/: compiled JavaScript, declarations, feature documentation, kernel sources, shared Zig source modules, and the corresponding optimized Wasm binaries. Development requires Zig 0.16, wasm-opt, and wasm-tools on PATH.

just memory-profile runs every owning data structure in an isolated Node process with explicit GC. It fails if live Wasm allocations do not return to baseline, allocator capacity keeps growing after warmup, or post-GC heap/external/ArrayBuffer memory does not plateau over the final rounds. Peak RSS is reported as an informational high-water mark because V8/Wasm tier-up and native allocators can retain code or arena pages after owned storage has been released.