npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@shift-labs/ferrosearch

v0.1.1

Published

In-memory full-text search engine for Bun and Node, implemented in Rust; MiniSearch-compatible

Downloads

1,584

Readme

ferrosearch

An in-memory full-text search engine for Bun and Node, implemented in Rust as a native addon.

ferrosearch provides exact, prefix, and fuzzy matching with BM25+ ranking, query combination trees, auto-suggestions, and a full document lifecycle. It is index- and API-compatible with MiniSearch 7: serialized indexes load in either library. A test suite of 14,000+ assertions verifies identical search results, scores, result ordering, and error messages.

Features

  • Exact match, prefix search, fuzzy match, field boosting, query combination trees (AND / OR / AND_NOT).
  • Auto-suggestion engine for query auto-completion.
  • BM25+ result ranking.
  • Add, remove, discard, and replace documents at any time.
  • Index serialization, interchangeable with MiniSearch in both directions.
  • JSON fast paths that keep bulk data out of the binding layer.
  • Result caching for repeated queries, invalidated on mutation.
  • Roughly 3x lower memory usage than an equivalent JavaScript index.

Installation

bun add @shift-labs/ferrosearch

Prebuilt binaries cover macOS (arm64, x64) and Linux (x64 and arm64, glibc and musl). Installation selects the matching binary automatically. On other platforms, build from source with Rust installed: bun install && bun run build.

import { FerroSearch } from "@shift-labs/ferrosearch";

For drop-in migration from MiniSearch, the same class is also exported under the alias MiniSearch.

Usage

Basic usage

const documents = [
  { id: 1, title: "Moby Dick", text: "Call me Ishmael. Some years ago...", category: "fiction" },
  { id: 2, title: "Zen and the Art of Motorcycle Maintenance", text: "I can see by my watch...", category: "fiction" },
  { id: 3, title: "Neuromancer", text: "The sky above the port was...", category: "fiction" },
  { id: 4, title: "Zen and the Art of Archery", text: "At first sight it must seem...", category: "non-fiction" },
  // ...and more
];

const index = new FerroSearch({
  fields: ["title", "text"], // fields to index for full-text search
  storeFields: ["title", "category"], // fields to return with search results
});

// Index all documents
index.addAll(documents);

// Search with default options
const results = index.search("zen art motorcycle");
// => [
//   { id: 2, title: 'Zen and the Art of Motorcycle Maintenance', category: 'fiction', score: 2.77258, match: { ... } },
//   { id: 4, title: 'Zen and the Art of Archery', category: 'non-fiction', score: 1.38629, match: { ... } }
// ]

Search options

// Search only specific fields
index.search("zen", { fields: ["title"] });

// Boost some fields (here "title")
index.search("zen", { boost: { title: 2 } });

// Prefix search ('moto' matches 'motorcycle')
index.search("moto", { prefix: true });

// Fuzzy search with a maximum edit distance of 0.2 * term length, rounded
// to the nearest integer ('ismael' matches 'ishmael')
index.search("ismael", { fuzzy: 0.2 });

// Combine terms with AND instead of the default OR
index.search("zen art", { combineWith: "AND" });

// Set default search options upon initialization
const tuned = new FerroSearch({
  fields: ["title", "text"],
  searchOptions: { boost: { title: 2 }, fuzzy: 0.2 },
});

Query combination trees

Combine subqueries with different operators and options:

// Documents that contain "zen" and ("motorcycle" or "archery")
index.search({
  combineWith: "AND",
  queries: [
    "zen",
    { combineWith: "OR", queries: ["motorcycle", "archery"] },
  ],
});

Wildcard search

The FerroSearch.wildcard symbol queries all documents; the equivalent wildcardSearch method is also available:

index.search(FerroSearch.wildcard);
index.wildcardSearch();

Auto suggestions

index.autoSuggest("zen ar");
// => [ { suggestion: 'zen archery art', terms: [ 'zen', 'archery', 'art' ], score: 1.73332 },
//      { suggestion: 'zen art', terms: [ 'zen', 'art' ], score: 1.21313 } ]

// Fuzzy suggestions for misspelled input
index.autoSuggest("neromancer", { fuzzy: 0.2 });
// => [ { suggestion: 'neuromancer', terms: [ 'neuromancer' ], score: 1.03998 } ]

ferrosearch ranks suggestions by the relevance of the documents that the suggested search would return.

Removing, discarding, and replacing documents

// Immediate removal: requires the full, unchanged document
index.remove(documents[0]);

// Discard by ID: faster, cleans up lazily via vacuuming
index.discard(2);

// Replace a document with a new version under the same ID
index.replace({ id: 3, title: "Neuromancer (2nd ed.)", text: "..." });

// Vacuuming removes lingering references to discarded documents. It runs
// automatically by default, or manually:
index.vacuum();

Serialization

const serialized = index.toJsonString();

// Later:
const restored = FerroSearch.loadJson(serialized, {
  fields: ["title", "text"],
  storeFields: ["title", "category"],
});

loadJson requires the same options that produced the serialized index. The serialization format is MiniSearch's version-2 format; indexes transfer between the two libraries in both directions.

JSON fast paths

Bulk data crosses the native boundary as JSON strings — once, instead of converting object graphs through the bindings. The package entry point already routes addAll and search through these paths; they are also available directly:

index.addAllJson(jsonArrayOfDocuments); // e.g. a file or network payload
const results = JSON.parse(index.searchJson("zen art", { prefix: true }));
const serialized = index.toJsonString();

Each fast path is exactly equivalent to its object-based counterpart. Documents, options, and queries must be JSON-serializable, which the native boundary requires in any case.

Result cache

ferrosearch memoizes search results per (query, options) and invalidates the cache on any mutation. Repeated queries — autocomplete keystrokes, dashboard refreshes — skip the engine. The cache only serves reads while no discards await vacuuming. In that state search is a pure function of the index, so cached results are observably identical to recomputed ones. Disable the cache with new FerroSearch({ ..., cache: false }).

MiniSearch compatibility

ferrosearch implements the behavior of minisearch 7.2.0. Given the same documents, options, and queries, it returns the same results: scores, result ordering (including ties), match data, serialized indexes, and error messages. The test suite verifies this equivalence against the minisearch package directly.

ferrosearch does not support function-valued options, because functions cannot cross the native boundary. This applies to tokenize, processTerm, extractField, stringifyField, filter, boostDocument, boostTerm, logger, and the function forms of prefix and fuzzy. The defaults run natively. Tokenization splits on Unicode space and punctuation, the engine lowercases terms, and field extraction reads plain object keys with JavaScript string coercion.

Equivalent alternatives:

  • Custom tokenization or stemming: pre-process documents into an indexed field before add, and pre-process queries before search.
  • filter: filter the returned results; the original applies filter to fully assembled results, so this is equivalent.
  • Nested field extraction: flatten the fields before indexing.

Additional differences:

  • A wildcard node inside a queries combination tree is not supported; use the wildcard query at the top level.
  • vacuum() is synchronous and complete; native code has no main thread to block. addAllAsync and loadJSONAsync are therefore not provided.
  • Index-corruption warnings (removing a changed document) go to stderr.
  • A partial bm25 object replaces the whole parameter set, as in the original. The resulting NaN scores surface as null — the value JSON.stringify produces — because JSON has no NaN.

Performance

The committed benchmarks run against minisearch 7.2.0 on the Billboard corpus (5,086 documents, two indexed fields). Measurements use Bun 1.3 on Apple Silicon, with warmup. Reproduce with bun run bench:compare and bun run bench:memory.

| Benchmark | Relative to minisearch | | --- | --- | | Serialize index (toJsonString) | 5x | | Load serialized index | 2.8x | | Auto suggestion | 1.6x | | Indexing from a JSON string (addAllJson) | 1.3x | | Indexing from JavaScript objects | 1.15x | | Selective queries (few results) | 1.0x | | Repeated queries (result cache) | 0.5x | | Exact / prefix / fuzzy search (cold, mixed incl. broad queries) | 0.3–0.5x |

Memory: approximately 6.8 MB RSS per index versus approximately 22 MB for the JavaScript implementation on the same corpus — about 3x smaller. The savings come from flat-vector posting lists, interned stored-field names, and enum-keyed document IDs.

Result transfer bounds bulk search performance: every result array crosses the native boundary as JSON. Parsing a thousand-result payload costs more than a warm JIT search in the JavaScript implementation. Bulk queries with large result sets therefore favor the JavaScript implementation. Indexing, serialization, loading, auto-suggestion, selective queries, and cold start favor ferrosearch.

Index loading streams the serialized index section directly into the engine without building an intermediate value tree, and posting lists are assembled in one pass. The loading advantage grows with index size: on a 12 MB knowledge-base index (5,000 long documents, ~15,000 terms), loading is 3.8x faster than the JavaScript implementation.

Implementation guarantees

  • Insertion-ordered structures replicate JavaScript Map iteration order throughout, so tie ordering and match-data ordering are identical to MiniSearch's.
  • ferrosearch measures edit distances in Unicode scalar values rather than UTF-16 code units; behavior differs only for astral-plane characters.
  • ferrosearch replicates the incremental index cleanup that MiniSearch performs during search after discard, including its order-dependent scoring bookkeeping.
  • The test suite (90 tests, 14,000+ assertions) compares every feature against the minisearch package: search batteries across corpora (unicode, mixed-type fields, edge-case IDs), full index-state equality after every lifecycle operation, error messages, deserialization of malformed and weird-but-valid indexes, and serialization round trips in both directions.

Development

bun install
bun run verify          # cargo fmt --check, clippy, cargo test, release build, bun test
bun run bench:compare   # speed benchmark
bun run bench:memory    # memory benchmark

The design document describes the internals: the slot-based radix tree, the reused-matrix fuzzy search, interned-term scoring, and the native-boundary strategy.

License

MIT. Derived from MiniSearch by Luca Ongaro (MIT).