npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@aiquants/fuzzy-search

v2.7.0

Published

Advanced fuzzy search library with Levenshtein distance, n-gram indexing, and Web Worker support

Readme

@aiquants/fuzzy-search

Advanced fuzzy search library with Levenshtein distance, n-gram indexing, and Web Worker support.

Features

  • Two-stage parallel search: an index stage (n-gram / word / phonetic candidate filtering, scored by the max of Jaro-Winkler similarity and a substring-containment score) and a Levenshtein distance stage run in parallel, and their results are merged by score
  • Web Worker offloading: both stages run in dedicated Web Workers shared globally per workerId, keeping the main thread responsive
  • Self-healing workers with a published status: the engine owns each worker's state (absent / loading / ready / failed), publishes it as an external store, terminates a failed worker at once, rejects the requests waiting on it immediately, and never fetches a URL that failed to load again (see Worker status and self-healing)
  • Japanese-aware: phonetic index (katakana → hiragana normalization), and index candidates work even for text shorter than the n-gram size — both a single-character query on its own and a single-character token inside a multi-token query (e.g. the 郎 in 佐藤 郎)
  • Substring-containment guarantee: when the query (case-normalized and trimmed of surrounding whitespace) is a literal substring of a field value, the index stage scores it at least containmentScoreBase (default 0.7) regardless of the Jaro-Winkler match window — a mid-string CJK match like "桜花" in "10 桜花" can no longer score 0 and be dropped by the threshold
  • Romaji / kana / width input (opt-in): with the normalize option both workers index normalized field values and read every query normalized, so yamada finds ヤマダ / やまだ / ヤマダ; the normalizer itself ships as the React-free, worker-free @aiquants/fuzzy-search/normalize entry (see Text normalization)
  • React integration: useFuzzySearch hook with debouncing, index pre-building, and search state management
  • TypeScript first: full type safety and IntelliSense support
  • Highly configurable: thresholds, sort modes, per-field weights, worker tuning options

Installation

# Using pnpm (recommended)
pnpm add @aiquants/fuzzy-search

# Using npm
npm install @aiquants/fuzzy-search

react / react-dom are optional peer dependencies — they are required only when you use the @aiquants/fuzzy-search/react entry.

⚠️ Breaking changes (v2)

  1. useFuzzySearch moved to the /react subpath. Import it from @aiquants/fuzzy-search/react. The main entry no longer depends on React, so React-less projects can use the core API.
  2. multiTermOperator now defaults to "or". The previous default "and" required every whitespace-separated token to appear as an exact substring, which effectively disabled fuzzy matching for multi-word queries. Pass multiTermOperator: "and" explicitly if you need the old narrowing behavior.
  3. CJS builds require explicit worker URLs. The CommonJS build only has a shim of import.meta.url (the file URL of dist/index.js under Node; the loading script's URL, or a guess relative to the document, in a browser), so its built-in default rarely points at the real worker files — pass workerUrls (see Worker setup). ESM builds resolve worker URLs automatically.
  4. The constructor now merges partial options with defaults. new FuzzySearchManager({ threshold: 0.5 }) works as expected (previously a partial object silently discarded all defaults and every search returned zero results).

Quick Start

import { FuzzySearchManager } from '@aiquants/fuzzy-search'

const items = [
  { name: 'Apple', category: 'Fruit' },
  { name: 'Banana', category: 'Fruit' },
  { name: 'Carrot', category: 'Vegetable' },
]

const searchManager = new FuzzySearchManager({ threshold: 0.4 })

const results = await searchManager.search('aple', items, ['name'])
// => [{ item: { name: 'Apple', ... }, score: ..., baseScore: ..., matchedFields: ['name'], originalIndex: 0 }]

results.forEach((result) => {
  console.log(`Item: ${result.item.name}, Score: ${result.score}`)
})

// Clean up (rejects in-flight searches and releases worker references)
searchManager.dispose()

A top-level threshold is propagated to both workers unless you specify indexWorkerOptions / levenshteinWorkerOptions explicitly.

Worker setup

The search runs in two Web Workers. How their scripts are located depends on your build setup:

  • ESM + bundler-free / modern bundlers: the default worker URLs are resolved relative to the library module via import.meta.url. No configuration needed.
  • Vite (recommended for apps): import the worker files as URLs and pass them explicitly, so they are emitted as build assets:
import indexWorkerUrl from '@aiquants/fuzzy-search/worker/indexWorker?url'
import levenshteinWorkerUrl from '@aiquants/fuzzy-search/worker/levenshteinWorker?url'

const manager = new FuzzySearchManager({
  workerUrls: {
    indexWorker: indexWorkerUrl,
    levenshteinWorker: levenshteinWorkerUrl,
  },
})
  • CommonJS (require): workerUrls is required. Without it the manager logs a warning and search() rejects because no worker can be created.
  • With the normalize option: both workers load a third script, the normalizer, with import() the first time a request carries the option (a page that never sets it never downloads it). With explicit worker URLs, pass its URL too:
import normalizerUrl from '@aiquants/fuzzy-search/worker/normalizer?url'

const manager = new FuzzySearchManager({
  normalize: { convertRomaji: true },
  workerUrls: { indexWorker: indexWorkerUrl, levenshteinWorker: levenshteinWorkerUrl, normalizer: normalizerUrl },
})

The workers load the normalizer from a URL the manager sends, never by a literal import(), so a bundler that re-bundles the worker scripts (Vite's ?worker&url, also in its default iife worker format) builds them unchanged; the normalizer is one self-contained module. Sizes as shipped (minified / gzip -9): indexWorker.mjs 22,012 / 6,278 B and levenshteinWorker.mjs 18,425 / 5,225 B, which every page with a manager downloads; normalizer.mjs 25,449 / 10,050 B, fetched only at the first request with the normalize option.

The worker URL rule

The engine resolves every worker URL with one rule, exported as FuzzySearchManager.resolveWorkerUrls(workerUrls?) so a host that needs the URL (to compare it with the published status, to log it) never re-implements it. Per script — indexWorker, levenshteinWorker, and normalizer (the script the workers load for the normalize option):

  1. the manager's workerUrls.<script> when it is a non-empty string, else
  2. the static FuzzySearchManager.defaultWorkerUrls.<script> when it is a non-empty string (read at the moment the engine creates the worker or sends a request), else
  3. the built-in default next to the library module (import.meta.url — in the CommonJS build only a shim of it, see above; null where no module URL exists at all).

A given string is resolved against the page origin (window.location.origin), and on a localhost / 127.0.0.1 page worker_file=true is added to the query (it keeps the Vite dev server from injecting HMR code into the worker). The function returns the absolute URLs exactly as the engine passes them to new Worker(...) — null for a worker whose URL cannot be resolved in the current environment (the default without a module URL, or a relative or unparsable URL without a page) — and never throws. Each (script, URL read, page origin and host name) is parsed once and remembered, so calling it at every render costs no URL parsing; the static defaultWorkerUrls is still read at every call, so a corrected static applies at the next call.

The members are one exported type, WorkerUrls ({ indexWorker?, levenshteinWorker?, normalizer? }), used by the workerUrls option, by FuzzySearchManager.defaultWorkerUrls and by every static that takes URLs; ResolvedWorkerUrls has the same members, each string | null.

FuzzySearchManager.resolveWorkerUrls({ indexWorker: '/assets/indexWorker.mjs' })
// => { indexWorker: 'https://app.example/assets/indexWorker.mjs', levenshteinWorker: '<built-in default>', normalizer: '<built-in default>' }

React Integration

import { useMemo } from 'react'
import { useFuzzySearch } from '@aiquants/fuzzy-search/react'

function SearchComponent({ items }: { items: User[] }) {
  // ✅ Memoize searchFields and options — see the note below
  const searchFields = useMemo(() => ['name', 'email'], [])
  const options = useMemo(() => ({ threshold: 0.4, debounceMs: 300 }), [])

  const {
    searchTerm,
    setSearchTerm,
    filteredItems,
    isSearching,
    isIndexBuilding,
    error,
  } = useFuzzySearch(items, searchFields, options)

  return (
    <div>
      <input
        value={searchTerm}
        onChange={(e) => setSearchTerm(e.target.value)}
        placeholder="Search..."
      />
      {isSearching && <div>Searching...</div>}
      <ul>
        {filteredItems.map((result) => (
          <li key={result.originalIndex}>{result.item.name}</li>
        ))}
      </ul>
      {error && <div role="alert">{error}</div>}
    </div>
  )
}

The hook also returns searchHistory, getSearchSuggestions, performanceMetrics, rebuildIndex, resetStats, clearSearch, and clearHistory. A separate useSearchHistory hook provides standalone history management.

Important: memoization required

useFuzzySearch compares the items array, searchFields, and options by reference. If you create new objects on every render, the hook re-initializes the search manager each time and can enter a render loop. Memoize them with useMemo (or define them outside the component).

In-place mutation of the same items array reference is not detected — pass a new array when the data changes (the standard React immutable-update pattern), or call rebuildIndex() explicitly.

Worker isolation vs sharing

The workerId option controls how Web Workers are managed:

  • Sharing (same ID, the default "default"): multiple FuzzySearchManager instances share the same workers and the built index is reused when the dataset is identical. Each request carries a dataset hash, so managers with different datasets never receive each other's results — but they will rebuild the shared index back and forth. For that case, prefer isolation.
  • Isolation (different IDs): separate worker instances per ID. Use this when multiple search components work on different datasets concurrently.
import { FuzzySearchManager, LogLevel } from '@aiquants/fuzzy-search'

const productSearch = new FuzzySearchManager({ workerId: 'products' })
const userSearch = new FuzzySearchManager({
  workerId: 'users',
  logLevel: LogLevel.DEBUG, // DEBUG / INFO / WARN / ERROR / NONE
  logger: console,          // optional custom ILogger implementation
})

Note: workers are kept alive after dispose() so they can be reused by later manager instances (they are page-scoped singletons per workerId). A live worker is shared whatever workerUrls a later manager of the same workerId passes; the URLs are read only when the engine has to create a worker.

Worker status and self-healing

The engine owns the state of each worker id's two workers and publishes it, so every manager of the id — whoever created it — reads the same state, and a host never has to mirror it from events. The registry lives in the engine module, and the package ships that module once: managers created through one installed copy of this package share it — new FuzzySearchManager(...) from the main entry and the useFuzzySearch hook of the /react entry alike (a second, separately bundled copy of the package, or the CommonJS and ESM builds loaded side by side, has its own registry and its own workers).

import { FuzzySearchManager } from '@aiquants/fuzzy-search'

const status = FuzzySearchManager.getWorkerStatus('products')   // workerId; undefined means "default"
status.index.state            // 'absent' | 'loading' | 'ready' | 'failed'
status.index.url              // the resolved URL of the live worker, or of the failed attempt
status.index.failure          // { phase: 'create' | 'load' | 'run', url, message } when failed, else null
status.index.rememberedFailures // every URL that failed to create or load under this id (never fetched again)
status.index.crashedDatasets  // every dataset this worker stopped on again when a search was re-sent (never sent to it again)
status.index.breaker          // the worker's circuit breaker: { open, probing, openedAt, nextProbeAt, failure, openings }, or null while closed

const unsubscribe = FuzzySearchManager.subscribeWorkerStatus('products', (next) => {
  console.log(next.levenshtein.state)
})
  • The snapshot is a store snapshot. getWorkerStatus() returns a frozen object whose identity changes only when that id's state changes (an unchanged worker keeps its object too), and listeners are called synchronously after every change. It can be passed straight to React's useSyncExternalStore, also as the server snapshot (an untouched id reads absent on both workers):

    const status = useSyncExternalStore(
      (onChange) => FuzzySearchManager.subscribeWorkerStatus(workerId, onChange),
      () => FuzzySearchManager.getWorkerStatus(workerId),
      () => FuzzySearchManager.getWorkerStatus(workerId),
    )
  • Whether a manager would get its workers is the engine's decision, asked without side effects. FuzzySearchManager.workerAcquisitionOf(workerId, workerUrls?) answers exactly when a manager of that worker id with those workerUrls gets both workers — the rule the acquisition itself follows, so a host never re-derives it from the status. Per worker (the index worker first, the Levenshtein worker only when the index worker can be had): the id holds it live (loading or ready, whatever its URL), or Web Workers exist and the resolved URL is not null and not remembered failed. It returns a frozen WorkerAcquisition:

    const acquisition = FuzzySearchManager.workerAcquisitionOf('products', { indexWorker: '/assets/indexWorker.mjs' })
    acquisition.allowed     // whether such a manager gets both workers
    acquisition.urls        // the resolved URLs (ResolvedWorkerUrls, frozen)
    acquisition.unresolved  // the workers whose URL cannot be resolved here ('index' first; 'levenshtein' only when the index worker can be had)
    acquisition.workers     // per worker, whether such a manager gets it: { index, levenshtein } (allowed is both)

    workers answers per worker by the same decision: index — the id holds it live, or it may be created; levenshtein — a live one always, a new one only when the index worker can be had. allowed is exactly workers.index && workers.levenshtein. Every call needs the index worker, so a host may hold a manager whenever workers.index is true: with workers.levenshtein false, searchWithStages() (below) answers with the index stage and reports the Levenshtein stage unavailable, while search() rejects as it always has.

    The query is pure: it records nothing in the status, notifies no one, logs nothing, creates and fetches nothing, and parses no URL for a repeated input. The answer stays the same object while its values are unchanged, so it (or its allowed) is a useSyncExternalStore snapshot with subscribeWorkerStatus; a host that wants to tell the user about an unresolvable URL reports unresolved itself:

    const acquisition = useSyncExternalStore(
      (onChange) => FuzzySearchManager.subscribeWorkerStatus(workerId, onChange),
      () => FuzzySearchManager.workerAcquisitionOf(workerId, workerUrls),
      () => null,
    )

    FuzzySearchManager.canAcquireWorkers(workerId, workerUrls?) is deprecated (removed in the next major version): it returns the same allowed but also records each unresolved worker in the id's status as the create failure with url: null that an acquisition records (subscribers hear it in a microtask).

  • States. absent → loading when a manager creates the worker; loading → ready at the worker's first message (the handshake every manager sends); failed with phase create (the Worker constructor threw, or no URL could be resolved), load (the worker's own error event fired before its first message, and no worker from that URL has answered under the id — the script could not be fetched or evaluated) or run (its own error event fired after it had answered, or before the first message of a worker re-created from a URL that has answered under the id: a URL that has answered once is proven and never becomes a remembered load failure).

  • A failed worker is terminated at once. The engine listens to every worker it creates (a failure is heard even when no manager is alive), terminates the failed worker, drops it, clears the shared index and dataset cache of the id, and settles every wait on it at once instead of letting it wait out its timeout: a load failure rejects them with a WorkerFailureError (worker, workerId, workerUrl, phase, and refusal: null for a death, 'dataset' or 'breaker' for a refusal described below), a run failure hands a waiting search to the re-send below. This includes the waits of a manager disposed meanwhile, so a search that joined its index build never waits out the build timeout.

  • A URL that failed to create or load is never fetched again for the page's life under that worker id and worker: managers whose URL is remembered get no such worker (their search() rejects as when the workers are unavailable), so a broken URL costs one fetch, not one per keystroke. Any other URL is a fresh attempt — the first manager (or request) with a different URL creates the worker from it.

  • A search survives a worker that stops after answering. When a worker dies with phase run while a search() waits on it, the engine re-creates the worker from that search's manager's URL and re-sends the search's part: the index part rebuilds the index on the new worker first, the Levenshtein part re-sends the dataset. A worker handles its requests one at a time in the order they were posted, so the engine blames the death on the dataset of the oldest request the worker had not answered (the one it was processing; an idle death blames none). A part of that dataset spends its one re-send per worker kind and search, so a search during which both workers die still resolves (each gets its own re-send); a part of any other dataset waiting on the same worker — another box's search under a shared worker id, or the previous options' search of the same box — is re-sent without spending it, so a dataset that never crashed a worker is never recorded below. Searches waiting together share the re-created worker and its rebuild. Every engine user gets this policy — direct search() callers and the useFuzzySearch hook alike (the hook does not turn a run-phase indexWorker:error into its error; a final failure reaches it through the rejected search).

  • A dataset a worker stops on twice is not sent to it again. When the same worker dies with phase run again during the re-send, the search rejects with a WorkerFailureError (phase run) whose message ends with (it stopped again when the request was re-sent to a new worker, so the engine sends this dataset to the <Index | Levenshtein> Worker no more; a changed dataset is sent once the worker's circuit breaker lets requests through), and the dataset is published in status.<worker>.crashedDatasets ({ dataHash, failure }, the list never shrinks). Any later search() or prebuildIndex() of that dataset that needs that worker rejects at once with the same message, posting nothing and creating no worker (the index worker is needed by every search, the Levenshtein worker when enableLevenshtein), so a deterministic crash (an out-of-memory index build, say) costs two crashes per dataset, not two per keystroke; a changed dataset is sent once the worker's circuit breaker (below) lets requests through, and a search with enableLevenshtein: false still gets the index stage's answer after a Levenshtein crash. The error's refusal is 'dataset' (on the final rejection too). rebuildIndex(), the explicit forced rebuild, is never refused. search() resolves with the full answer of every enabled stage or rejects; it never resolves with one stage's rows only. searchWithStages() (below) is the same search that answers with the index stage when only the Levenshtein worker cannot serve, and says so.

  • A circuit breaker per worker bounds content that keeps crashing it. A crash caused by content every new dataset keeps (a pathological row in an append stream, a size limit) would otherwise cost two crashes, two re-created workers and two whole-dataset posts per dataset change. So the final crash that records a dataset also opens that worker's breaker under the worker id (status.<worker>.breaker, frozen):

    status.levenshtein.breaker
    // { open: true, probing: false, openedAt: 1767225600000, nextProbeAt: 1767225605000,
    //   failure: { phase: 'run', url, message }, openings: 1 }

    While open is true, every search() needing that worker and every prebuildIndex() (index worker) rejects at once — before any worker is created or anything is posted — with a WorkerFailureError (phase run, refusal 'breaker') whose message ends with (the engine's circuit breaker for the <Index | Levenshtein> Worker of this workerId opened after this failure (opening <n>): it refuses every request until <ISO time>, then lets one request through to probe it). The window is 5 s for the first opening and doubles with every consecutive opening (10 s, 20 s, 40 s, 80 s, 160 s), capped at 300 s; nextProbeAt is when it ends (epoch milliseconds, like openedAt). At nextProbeAt the engine publishes open: false, and the next call needing that worker is let through as the probe (open and probing are true while it is in flight, and other calls are refused with … it refuses every request while the one request it let through to probe it is answered)). probing is true exactly while the probe is in flight, so open && !probing is the window itself, and a host that reads probing waits for the probe's outcome — the next status publication — instead of treating its own pending call as refused. When the probe resolves the breaker closes (breaker becomes null, and the next opening starts from 5 s again); when the probe's own request kills the worker (a search after its re-send, a prebuildIndex() at its one death) the breaker opens again with the next window; when the probe ends any other way the next call is the probe. So such content costs at most one opening (two crashes) per window, however often the dataset changes. Only the workers a call needs are checked: with only the Levenshtein breaker open, a search with enableLevenshtein: false gets the index stage's answer. Calls already running when the breaker opens finish their parts; rebuildIndex() is never refused and does not touch the breaker. crashedDatasets keeps listing only the datasets whose re-send died. A prebuildIndex() death while the breaker is closed opens nothing. A refused call reads its dataset only when a dataset with the same item count and search fields is recorded for that worker (equal datasets have equal shapes), so the calls of an append stream refused during a window hash nothing; a refusal of prebuildIndex() is logged at debug level (a refused search() is not logged; its caller gets the rejection).

  • A detailed search reports what ran. searchWithStages(query, items, searchFields, options?) is search() with the Levenshtein worker optional — one implementation, so it runs, re-sends and rejects exactly as search() does except when only the Levenshtein worker cannot serve the call. The engine's admission decides per worker: the index worker is required (its refusal or failure rejects as search() does), and when the Levenshtein worker's breaker refuses the dataset, the dataset is recorded for it, it cannot be had, its script fails to load, or the call's re-send dies on it again, the call resolves with the index stage's answer:

    const { results, stages } = await manager.searchWithStages('aple', items, ['name'])
    stages.index        // { ran: true }
    stages.levenshtein  // { ran: true } or { ran: false, reason: 'breaker' | 'dataset' | 'crashed' | 'unavailable' | 'disabled' | … }

    Reasons (frozen reports): blank (blank query; both stages), disabled (the stage is off in the options), disposed (the manager was disposed first; results []), superseded / contended (index stage: a newer dataset of this manager overtook the search, or the shared index still held another dataset after one rebuild), breaker (the Levenshtein breaker refused the dataset at admission — its window, or another call is its probe), dataset (recorded for the Levenshtein worker before the call), crashed (recorded during the call) and unavailable (the Levenshtein worker could not be had, or its script failed to load). blank, disabled, dataset, crashed and unavailable are permanent for the manager, the dataset and the options; breaker, superseded and contended are temporary — an answer carrying one of them is not the dataset's complete answer and should be searched again later. A Levenshtein failure outside that list (its 30 s timeout, an error reply) rejects as in search(). A call that resolves closes the breaker probe of each worker whose stage served it and ends the probe of a Levenshtein worker whose stage could not serve (that breaker re-opens on a death blamed on the call's dataset, otherwise the next call probes); it emits the same search:* events and updates the same stats.

  • Outside a search, a worker that failed after answering is recreated by the next request that uses it. Nothing is remembered for a run failure: the next search() / prebuildIndex() / rebuildIndex() of any manager of the id that uses the worker and is not refused creates it again from that manager's resolved URL (and rebuilds the index on it: prebuildIndex() skips only when the id's shared index already holds the data). prebuildIndex() and rebuildIndex() re-send nothing: a worker death during them rejects them. Every request first re-acquires, of the workers it uses, the ones its manager lacks, not only the constructor: a search uses the index worker and, with enableLevenshtein, the Levenshtein worker; prebuildIndex() and rebuildIndex() use the index worker alone; a search part's re-send re-acquires only the worker that died. A worker a request does not use is never created for it, so a search with enableLevenshtein: false neither re-creates a stopped Levenshtein worker nor needs one. The constructor prepares both workers.

  • Events stay as they were, with more in them. workers:initialized (once per manager, at the end of its constructor's setup) keeps indexWorker / levenshteinWorker and adds workerId, workerUrls (the manager's resolved URLs) and status. For a worker's own error event, indexWorker:error / levenshteinWorker:error still deliver the original Event, now carrying worker, workerId, workerUrl and phase; an error reply to one request with no waiter is still the { requestId, error, processingTime } record.

Text normalization

The engine compares raw text unless you pass normalize. With it, both workers index normalizeText(String(item[field] ?? ""), normalize) for every search field and read every query as normalizeText(query, normalize); all matching (candidates, Jaro-Winkler, containment, Levenshtein, the "and" token check) then runs on those texts. Results still carry the original items, originalIndex and matchedFields (names of the original fields).

import { FuzzySearchManager } from "@aiquants/fuzzy-search"

const manager = new FuzzySearchManager<{ label: string }>({ normalize: { convertRomaji: true } })
const items = [{ label: "ヤマダ" }, { label: "やまだ" }, { label: "ヤマダ" }, { label: "山田" }]
const results = await manager.search("yamada", items, ["label"])
// ヤマダ, やまだ and ヤマダ are found (each normalizes to ヤマダ); 山田 is not (kanji have no reading here)
  • Absent (undefined): no normalization — the engine behaves exactly as 2.2.x (same matching, same index keys).
  • An object: omitted flags take DEFAULT_NORMALIZE_OPTIONS (convertCase, convertWidth, convertKana, normalizeVariants on; convertRomaji, ignoreSpaces, ignoreSymbols, ignoreLongVowels off), so {} already makes hiragana ↔ katakana and half-width ↔ full-width equal. It is accepted by the constructor, by the per-call options of search() / prebuildIndex() / rebuildIndex() (a call's own normalize key wins, normalize: undefined turns it off for that call) and by useFuzzySearch.
  • Validated on the main thread by resolveNormalizeOptions(value, "[fuzzy-search] normalize"): the constructor throws, and search() (even for a blank query) / prebuildIndex() / rebuildIndex() reject, with a RangeError before any worker is touched — [fuzzy-search] normalize must be a plain object; got 5, [fuzzy-search] normalize keys must be one of "convertCase", "convertWidth", "convertKana", "convertRomaji", "normalizeVariants", "ignoreSpaces", "ignoreSymbols", "ignoreLongVowels"; got "convertCaps", [fuzzy-search] normalize.convertCase must be a boolean; got 1, [fuzzy-search] normalize.convertKana must be true when convertRomaji is true; got false. The resolved options travel to the workers inside the request options (structured clone).
  • A query that normalizes to nothing ("!?" under ignoreSymbols) returns no results.
  • Shared workers: the dataset key includes the resolved flags, so managers of one workerId with different (or no) normalization never reuse each other's index.
  • Case: normalize.convertCase folds case first; caseSensitive then applies to the normalized texts.
  • Positions: the engine reports no character positions, only matchedFields (names of the original fields). The position-mapped forms that map a range of a normalized form back to the source text (toMappedFormsInto, MappedText) are part of the normalize entry's query compilation forms.

Equivalences (each line run through the real normalizeText):

| Input | Options | Normalized | | --- | --- | --- | | yamada, YAMADA, やまだ, ヤマダ, ヤマダ | { convertRomaji: true } | ヤマダ | | Tōkyō | { convertRomaji: true } | トーキョー | | toukyou, トウキョウ | { convertRomaji: true } | トウキョウ | | ko-nsuta-chi, コーンスターチ | { convertRomaji: true } | コーンスターチ | | Matuyama | { convertRomaji: true } | マツヤマ | | ヤマダ, ヤマダ | {} | ヤマダ | | ABC | {} | abc | | Straße | {} | strasse | | ポテト-25kg, ぽてとー25kg | {} | ポテトー25kg | | ㄱㅏ, 가 | {} | 가 | | コーン, コオン | { ignoreLongVowels: true } | コン | | Shinʼya | { ignoreSymbols: true } | shinya | | 山田 太郎 | { ignoreSpaces: true } | 山田太郎 |

Limits:

  • Kanji readings are out of scope: 山田 stays 山田. To find kanji by reading, put the reading in the data (a reading field such as { label: "山田", reading: "やまだ" }) and search both fields.
  • One reading per text: fields and queries are read alike, with the option-text reading of romaji (ti / tu / di / du / wo as チ / ツ / ヂ / ヅ / ヲ, so tisshu reads チッシュ). The alternative query readings a component may compile (ティ for ti) are not applied by the engine.
  • A normalized text is a whole-text key: normalization reads characters from their neighbours, so a fragment can read differently inside the word (under convertRomaji, ation reads アチオン while nation reads ナチオン; abc reads アbc). Such fragments still meet through the fuzzy stages, but not as literal substrings.

Cost: the normalizer runs inside each worker, at index build (every field value once) and once per query. The worker scripts do not carry it: each worker loads worker/normalizer.mjs (25.4 KB, 10.1 KB gzip) once, at the first request with the option, from the URL the manager resolves (see Worker setup); a request whose normalizer cannot be loaded rejects with [fuzzy-search] the normalizer could not be loaded from <url>: …, and search() / prebuildIndex() / rebuildIndex() with the option reject before touching a worker when no normalizer URL can be resolved ([fuzzy-search] normalize: no URL could be resolved for the normalizer script (specify options.workerUrls.normalizer explicitly)). The main entry validates the option with the ./normalize entry's options module, which it shares in the ESM build.

The normalize entry (@aiquants/fuzzy-search/normalize)

The same normalizer, importable on its own: no React and no worker in its import graph, ESM split per module so a bundler keeps only what is used (measured with esbuild, minified: hiraganaToKatakana alone 2.6 KB, resolveNormalizeOptions 1.4 KB, romajiToKatakana 8.5 KB, normalizeText 25.4 KB, getNormalizedVariants 30.6 KB, the whole entry 43.7 KB).

import { getNormalizedVariants, hiraganaToKatakana, katakanaToHiragana, normalizeText, romajiToKatakana } from "@aiquants/fuzzy-search/normalize"

normalizeText("ヤマダ") // "ヤマダ"
normalizeText("yamada", { convertRomaji: true }) // "ヤマダ"
romajiToKatakana("yamada") // "ヤマダ"
hiraganaToKatakana("やまだ") // "ヤマダ"
katakanaToHiragana("ヤマダ") // "やまだ"
getNormalizedVariants("ヤマダ") // ["ヤマダ", "ヤマダ", "やまだ"]

Documented API: normalizeText(text, options?), getNormalizedVariants(text, options?), normalizeJapaneseVariants(text), hiraganaToKatakana(text), katakanaToHiragana(text), romajiToKatakana(text), resolveNormalizeOptions(options, name) (name is the setting as the caller names it, owner first, and starts every message — the engine passes [fuzzy-search] normalize, a library validating its own setting passes e.g. [my-lib] normalizeOptions), DEFAULT_NORMALIZE_OPTIONS, and the types NormalizeOptions / ResolvedNormalizeOptions. Every function rejects a non-string text with a RangeError, and normalizeText / getNormalizedVariants / resolveNormalizeOptions reject invalid options the same way ([fuzzy-search] normalizeText: text must be a string; got 42, [fuzzy-search] normalizeText: options.convertCase must be a boolean; got 1).

Under moduleResolution: node10 (which ignores the exports map) the normalize and react subpaths resolve their declarations through the package's typesVersions.

Query compilation forms

The entry also exports the forms a host compiles its own queries over — documented API, covered by semver like the functions above (a change to any of them is a major release):

| Group | Names | | --- | --- | | Comparison forms of a text | FORM_TIERS, toTextForms, toTextFormsInto, toMappedFormsInto, InteriorPairs, interiorEndAt, distinctForms, trimmedRange; types FormSet, FormKind, FormTier, TextForms, MappedForms, MappedText, PartialReadings, InteriorCursor, DistinctForm | | Query reading and needles | readQuery, segmentQuery, isBlankQuery, queryNeedles, occurrenceTier, isAnswerTier; types QuerySegmentation, QueryNeedle, Tier | | Per-engine readings | ENGINE_READING_CONVENTIONS, engineForms, engineFieldsOf, engineWordAnchorsOf, placeBesideOthers, perEngineReading, LONE_WORD_PLACE; types EngineForms, EngineReadingConvention, EngineReadings, EngineWordAnchors, QueryWordPlace | | Grapheme clusters | isClusterBoundary, clusterCountOf, holdsMultiUnitCluster |

Each name's contract — its inputs and outputs, what each tier means, the order FORM_TIERS guarantees, what typing changes, the reading-interior protocol of InteriorPairs, the shapes of the Hangul IME stand-ins and the status of the jamo members — is in Query compilation forms, shipped in the package as docs/query-compilation-forms.md.

While a query is being typed (queryNeedles(query, options, true)), the needles also hold tier-2 Hangul IME stand-ins: what a Korean two-set IME can still turn the syllable it is composing into, each an answer-tier reading (literal, canonical, the long-vowel fold and each romaji reading) of a real continuation of the query, lone compound consonants included (ㄱ also stands for ㄳ, ㄳ for ㄱ + ᄉ), so a separator ignoreSpaces / ignoreSymbols removes never lets a compound compose onto the syllable it committed (가 ㄱ stands in as 가 ᆪ, not 갃). They are jamo texts, compared with the jamo image of every form a host stores; the page gives their shapes and how to match them in composed forms. FormSet.literalJamo / canonicalJamo (the images of the literal and canonical forms) are own data properties like every other member. distinctForms(forms, textOf, tiers) reads only the kinds in tiers, which must be FORM_TIERS or a subsequence of it (else a RangeError naming the first offending entry).

Nothing else is exported: the entry exports 51 names (the 10 documented normalization names, and the 24 values and 17 types of this table), and its export list is pinned by a test, so it cannot grow or shrink silently.

Configuration Options

Top-level options (FuzzySearchOptions, all optional in the constructor / hook):

| Option | Default | Description | | --- | --- | --- | | threshold | 0.4 (Index) / 0.3 (Levenshtein) | Minimum similarity score (0-1). Propagated to worker options if not explicitly overridden | | caseSensitive | false | Case-sensitive matching | | multiTermOperator | "or" | How whitespace-separated query tokens combine. "and" requires every token as an exact substring | | ngramSize | 2 | N-gram size for the index. Shorter queries fall back to substring matching over indexed keys | | minNgramOverlap | 1 | Per-token lower bound of the n-gram overlap threshold in the accurate candidate strategy | | sortBy | "relevance" | "score" | "relevance" (score + bonuses for fields meeting the threshold and actual partial matches) | "original" (input order) | | sortOrder | "desc" | Sort direction. Not applied when sortBy is "original" (always input order) | | enableIndexFiltering | true | Run the index (candidate) stage | | enableLevenshtein | true | Run the Levenshtein stage (full scan) | | parallelSearchStrategy | "balanced" | Result merging: "balanced" (merge by score) | "index-first" | "levenshtein-first" | | customWeights | {} | Per-field score multipliers, e.g. { name: 2.0 } | | debounceMs | 300 | Debounce for the React hook | | autoSearchOnIndexRebuild | true | Re-run the current query after the hook rebuilds the index | | workerId | "default" | Worker sharing/isolation key | | workerUrls | – | Explicit script URLs { indexWorker, levenshteinWorker, normalizer } (required for CJS; normalizer only with normalize); resolved by the worker URL rule | | logLevel / logger | WARN / console | Logging control. The logger object stays on the main thread (it is not sent to workers) | | normalize | – | Text normalization of every field value and query in both workers (see Text normalization); absent = raw text, as 2.2.x |

Worker tuning options: indexWorkerOptions (strategy: "fast" | "accurate" | "hybrid", threshold, ngramOverlapThreshold, minCandidatesRatio, maxCandidatesRatio, jaroWinklerPrefix, maxResults, relevanceFieldWeight, relevancePerfectMatchBonus, containmentScoreBase, containmentCoverageWeight) and levenshteinWorkerOptions (threshold, lengthSimilarityThreshold, lengthDiffPenalty, partialMatchBonus, maxResults, relevanceFieldWeight). Explicit 0 values are honored.

The index stage scores each field as max(JaroWinkler, containment): when the case-normalized, trimmed query is a literal substring of the field value, the containment score containmentScoreBase + containmentCoverageWeight * (trimmedQueryLength / valueLength) (defaults 0.7 / 0.3; an identical string scores exactly 1.0) rescues matches the Jaro-Winkler window cannot reach.

The scoring query is trimmed once up front, so Jaro-Winkler, containment, and the perfect-match equality check are all invariant to surrounding whitespace. In sortBy: "relevance" mode, the perfect-match bonus (relevancePerfectMatchBonus) is granted only on explicit string equality — not inferred from a high score.

For queries with more than one token, each field is scored against the whole query and against each token individually, and the best of those wins. Without this, tokens spread across different fields (e.g. "佐藤 商事" where one field holds "佐藤 商" and another "商事ファシリティーズ") leave every field failing to contain the whole query, so the containment bonus never applies and the whole query's length dilutes the similarity. The relevance bonus follows the same rule: a field matching any single token qualifies — by string equality in the index stage (relevancePerfectMatchBonus), by substring containment in the Levenshtein stage (partialMatchBonus).

Candidate retrieval treats whitespace-separated query tokens with OR semantics at every stage: the fast strategy collects n-grams from the whole query and from each token individually, the candidate-overflow filter and the accurate strategy evaluate n-gram overlap per token (an item survives if any one token clears its threshold), and tokens shorter than ngramSize — which cannot produce n-grams — fall back to substring matching over indexed keys instead of being silently dropped. A mixed query like "A 粉末" therefore keeps the items that only match the short token, and "佐藤 郎" returns the union of both tokens' hits rather than dropping the single-character token's side entirely (all three strategies return the same union).

API Reference

Main entry (@aiquants/fuzzy-search)

  • FuzzySearchManager<T> — search(), searchWithStages(), prebuildIndex(), rebuildIndex(), getStats(), getCurrentStats(), resetStats(), getSearchHistory(), updateSearchHistory(), on() / off() (search lifecycle events), dispose()
  • Statics: FuzzySearchManager.resolveWorkerUrls(workerUrls?), FuzzySearchManager.workerAcquisitionOf(workerId, workerUrls?), FuzzySearchManager.canAcquireWorkers(workerId, workerUrls?) (deprecated), FuzzySearchManager.getWorkerStatus(workerId?), FuzzySearchManager.subscribeWorkerStatus(workerId, listener), FuzzySearchManager.defaultWorkerUrls
  • WorkerFailureError, DEFAULT_FUZZY_SEARCH_OPTIONS, isSearchResultItem, LogLevel
  • Types: FuzzySearchOptions, PartialFuzzySearchOptions, SearchResultItem, SearchResults, StagedSearchResults, SearchStages, SearchStageReport, SearchStageSkipReason, SearchStats, UseFuzzySearchOptions, ILogger, WorkerStatus, WorkerSlotStatus, WorkerState, WorkerFailure, WorkerFailurePhase, CrashedDataset, WorkerBreaker, WorkerRefusal, WorkerUrls, ResolvedWorkerUrls, WorkerAcquisition, EngineWorkerKind, WorkerErrorEvent, WorkerErrorDetails, WorkersInitializedEvent, …

React entry (@aiquants/fuzzy-search/react)

  • useFuzzySearch(items, searchFields, options?) — main search hook
  • useSearchHistory(maxHistory?) — standalone search-term history

Normalize entry (@aiquants/fuzzy-search/normalize)

  • normalizeText, getNormalizedVariants, normalizeJapaneseVariants, hiraganaToKatakana, katakanaToHiragana, romajiToKatakana, resolveNormalizeOptions, DEFAULT_NORMALIZE_OPTIONS
  • Types: NormalizeOptions, ResolvedNormalizeOptions
  • The query compilation forms (see Query compilation forms)

Worker entries

  • @aiquants/fuzzy-search/worker/indexWorker
  • @aiquants/fuzzy-search/worker/levenshteinWorker
  • @aiquants/fuzzy-search/worker/normalizer (loaded by both workers for the normalize option)

Performance Tips

  1. Pre-build the index with prebuildIndex() (the React hook does this automatically when items changes)
  2. Keep the items array reference stable between searches — the dataset content hash is memoized per array reference, so a repeated call with the same array hashes nothing (mutating that array in place is not detected: pass a new array or call rebuildIndex())
  3. Use separate workerIds for components searching different datasets
  4. Set maxResults in the worker options for very large result sets
  5. Adjust threshold — lower values return more (and noisier) results

Browser Support

  • Chrome / Chromium 60+
  • Firefox 55+
  • Safari 11+
  • Edge 79+

License

MIT License

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.