@aiquants/fuzzy-search
v2.7.0
Published
Advanced fuzzy search library with Levenshtein distance, n-gram indexing, and Web Worker support
Readme
@aiquants/fuzzy-search
Advanced fuzzy search library with Levenshtein distance, n-gram indexing, and Web Worker support.
Features
- Two-stage parallel search: an index stage (n-gram / word / phonetic candidate filtering, scored by the max of Jaro-Winkler similarity and a substring-containment score) and a Levenshtein distance stage run in parallel, and their results are merged by score
- Web Worker offloading: both stages run in dedicated Web Workers shared globally per
workerId, keeping the main thread responsive - Self-healing workers with a published status: the engine owns each worker's state (absent / loading / ready / failed), publishes it as an external store, terminates a failed worker at once, rejects the requests waiting on it immediately, and never fetches a URL that failed to load again (see Worker status and self-healing)
- Japanese-aware: phonetic index (katakana → hiragana normalization), and index candidates work even for text shorter than the n-gram size — both a single-character query on its own and a single-character token inside a multi-token query (e.g. the
郎in佐藤 郎) - Substring-containment guarantee: when the query (case-normalized and trimmed of surrounding whitespace) is a literal substring of a field value, the index stage scores it at least
containmentScoreBase(default0.7) regardless of the Jaro-Winkler match window — a mid-string CJK match like"桜花"in"10 桜花"can no longer score 0 and be dropped by the threshold - Romaji / kana / width input (opt-in): with the
normalizeoption both workers index normalized field values and read every query normalized, soyamadafindsヤマダ/やまだ/ヤマダ; the normalizer itself ships as the React-free, worker-free@aiquants/fuzzy-search/normalizeentry (see Text normalization) - React integration:
useFuzzySearchhook with debouncing, index pre-building, and search state management - TypeScript first: full type safety and IntelliSense support
- Highly configurable: thresholds, sort modes, per-field weights, worker tuning options
Installation
# Using pnpm (recommended)
pnpm add @aiquants/fuzzy-search
# Using npm
npm install @aiquants/fuzzy-searchreact / react-dom are optional peer dependencies — they are required only when you use the @aiquants/fuzzy-search/react entry.
⚠️ Breaking changes (v2)
useFuzzySearchmoved to the/reactsubpath. Import it from@aiquants/fuzzy-search/react. The main entry no longer depends on React, so React-less projects can use the core API.multiTermOperatornow defaults to"or". The previous default"and"required every whitespace-separated token to appear as an exact substring, which effectively disabled fuzzy matching for multi-word queries. PassmultiTermOperator: "and"explicitly if you need the old narrowing behavior.- CJS builds require explicit worker URLs. The CommonJS build only has a shim of
import.meta.url(the file URL ofdist/index.jsunder Node; the loading script's URL, or a guess relative to the document, in a browser), so its built-in default rarely points at the real worker files — passworkerUrls(see Worker setup). ESM builds resolve worker URLs automatically. - The constructor now merges partial options with defaults.
new FuzzySearchManager({ threshold: 0.5 })works as expected (previously a partial object silently discarded all defaults and every search returned zero results).
Quick Start
import { FuzzySearchManager } from '@aiquants/fuzzy-search'
const items = [
{ name: 'Apple', category: 'Fruit' },
{ name: 'Banana', category: 'Fruit' },
{ name: 'Carrot', category: 'Vegetable' },
]
const searchManager = new FuzzySearchManager({ threshold: 0.4 })
const results = await searchManager.search('aple', items, ['name'])
// => [{ item: { name: 'Apple', ... }, score: ..., baseScore: ..., matchedFields: ['name'], originalIndex: 0 }]
results.forEach((result) => {
console.log(`Item: ${result.item.name}, Score: ${result.score}`)
})
// Clean up (rejects in-flight searches and releases worker references)
searchManager.dispose()A top-level threshold is propagated to both workers unless you specify indexWorkerOptions / levenshteinWorkerOptions explicitly.
Worker setup
The search runs in two Web Workers. How their scripts are located depends on your build setup:
- ESM + bundler-free / modern bundlers: the default worker URLs are resolved relative to the library module via
import.meta.url. No configuration needed. - Vite (recommended for apps): import the worker files as URLs and pass them explicitly, so they are emitted as build assets:
import indexWorkerUrl from '@aiquants/fuzzy-search/worker/indexWorker?url'
import levenshteinWorkerUrl from '@aiquants/fuzzy-search/worker/levenshteinWorker?url'
const manager = new FuzzySearchManager({
workerUrls: {
indexWorker: indexWorkerUrl,
levenshteinWorker: levenshteinWorkerUrl,
},
})- CommonJS (
require):workerUrlsis required. Without it the manager logs a warning andsearch()rejects because no worker can be created. - With the
normalizeoption: both workers load a third script, the normalizer, withimport()the first time a request carries the option (a page that never sets it never downloads it). With explicit worker URLs, pass its URL too:
import normalizerUrl from '@aiquants/fuzzy-search/worker/normalizer?url'
const manager = new FuzzySearchManager({
normalize: { convertRomaji: true },
workerUrls: { indexWorker: indexWorkerUrl, levenshteinWorker: levenshteinWorkerUrl, normalizer: normalizerUrl },
})The workers load the normalizer from a URL the manager sends, never by a literal import(), so a bundler that re-bundles the worker scripts (Vite's ?worker&url, also in its default iife worker format) builds them unchanged; the normalizer is one self-contained module.
Sizes as shipped (minified / gzip -9): indexWorker.mjs 22,012 / 6,278 B and levenshteinWorker.mjs 18,425 / 5,225 B, which every page with a manager downloads; normalizer.mjs 25,449 / 10,050 B, fetched only at the first request with the normalize option.
The worker URL rule
The engine resolves every worker URL with one rule, exported as FuzzySearchManager.resolveWorkerUrls(workerUrls?) so a host that needs the URL (to compare it with the published status, to log it) never re-implements it. Per script — indexWorker, levenshteinWorker, and normalizer (the script the workers load for the normalize option):
- the manager's
workerUrls.<script>when it is a non-empty string, else - the static
FuzzySearchManager.defaultWorkerUrls.<script>when it is a non-empty string (read at the moment the engine creates the worker or sends a request), else - the built-in default next to the library module (
import.meta.url— in the CommonJS build only a shim of it, see above;nullwhere no module URL exists at all).
A given string is resolved against the page origin (window.location.origin), and on a localhost / 127.0.0.1 page
worker_file=true is added to the query (it keeps the Vite dev server from injecting HMR code into the worker). The
function returns the absolute URLs exactly as the engine passes them to new Worker(...) — null for a worker whose
URL cannot be resolved in the current environment (the default without a module URL, or a relative or unparsable URL
without a page) — and never throws. Each (script, URL read, page origin and host name) is parsed once and remembered,
so calling it at every render costs no URL parsing; the static defaultWorkerUrls is still read at every call, so a
corrected static applies at the next call.
The members are one exported type, WorkerUrls ({ indexWorker?, levenshteinWorker?, normalizer? }), used by the workerUrls option, by FuzzySearchManager.defaultWorkerUrls and by every static that takes URLs; ResolvedWorkerUrls has the same members, each string | null.
FuzzySearchManager.resolveWorkerUrls({ indexWorker: '/assets/indexWorker.mjs' })
// => { indexWorker: 'https://app.example/assets/indexWorker.mjs', levenshteinWorker: '<built-in default>', normalizer: '<built-in default>' }React Integration
import { useMemo } from 'react'
import { useFuzzySearch } from '@aiquants/fuzzy-search/react'
function SearchComponent({ items }: { items: User[] }) {
// ✅ Memoize searchFields and options — see the note below
const searchFields = useMemo(() => ['name', 'email'], [])
const options = useMemo(() => ({ threshold: 0.4, debounceMs: 300 }), [])
const {
searchTerm,
setSearchTerm,
filteredItems,
isSearching,
isIndexBuilding,
error,
} = useFuzzySearch(items, searchFields, options)
return (
<div>
<input
value={searchTerm}
onChange={(e) => setSearchTerm(e.target.value)}
placeholder="Search..."
/>
{isSearching && <div>Searching...</div>}
<ul>
{filteredItems.map((result) => (
<li key={result.originalIndex}>{result.item.name}</li>
))}
</ul>
{error && <div role="alert">{error}</div>}
</div>
)
}The hook also returns searchHistory, getSearchSuggestions, performanceMetrics, rebuildIndex, resetStats, clearSearch, and clearHistory. A separate useSearchHistory hook provides standalone history management.
Important: memoization required
useFuzzySearch compares the items array, searchFields, and options by reference. If you create new objects on every render, the hook re-initializes the search manager each time and can enter a render loop. Memoize them with useMemo (or define them outside the component).
In-place mutation of the same items array reference is not detected — pass a new array when the data changes (the standard React immutable-update pattern), or call rebuildIndex() explicitly.
Worker isolation vs sharing
The workerId option controls how Web Workers are managed:
- Sharing (same ID, the default
"default"): multipleFuzzySearchManagerinstances share the same workers and the built index is reused when the dataset is identical. Each request carries a dataset hash, so managers with different datasets never receive each other's results — but they will rebuild the shared index back and forth. For that case, prefer isolation. - Isolation (different IDs): separate worker instances per ID. Use this when multiple search components work on different datasets concurrently.
import { FuzzySearchManager, LogLevel } from '@aiquants/fuzzy-search'
const productSearch = new FuzzySearchManager({ workerId: 'products' })
const userSearch = new FuzzySearchManager({
workerId: 'users',
logLevel: LogLevel.DEBUG, // DEBUG / INFO / WARN / ERROR / NONE
logger: console, // optional custom ILogger implementation
})Note: workers are kept alive after dispose() so they can be reused by later manager instances (they are page-scoped singletons per workerId). A live worker is shared whatever workerUrls a later manager of the same workerId passes; the URLs are read only when the engine has to create a worker.
Worker status and self-healing
The engine owns the state of each worker id's two workers and publishes it, so every manager of the id — whoever created it — reads the same state, and a host never has to mirror it from events. The registry lives in the engine module, and the package ships that module once: managers created through one installed copy of this package share it — new FuzzySearchManager(...) from the main entry and the useFuzzySearch hook of the /react entry alike (a second, separately bundled copy of the package, or the CommonJS and ESM builds loaded side by side, has its own registry and its own workers).
import { FuzzySearchManager } from '@aiquants/fuzzy-search'
const status = FuzzySearchManager.getWorkerStatus('products') // workerId; undefined means "default"
status.index.state // 'absent' | 'loading' | 'ready' | 'failed'
status.index.url // the resolved URL of the live worker, or of the failed attempt
status.index.failure // { phase: 'create' | 'load' | 'run', url, message } when failed, else null
status.index.rememberedFailures // every URL that failed to create or load under this id (never fetched again)
status.index.crashedDatasets // every dataset this worker stopped on again when a search was re-sent (never sent to it again)
status.index.breaker // the worker's circuit breaker: { open, probing, openedAt, nextProbeAt, failure, openings }, or null while closed
const unsubscribe = FuzzySearchManager.subscribeWorkerStatus('products', (next) => {
console.log(next.levenshtein.state)
})The snapshot is a store snapshot.
getWorkerStatus()returns a frozen object whose identity changes only when that id's state changes (an unchanged worker keeps its object too), and listeners are called synchronously after every change. It can be passed straight to React'suseSyncExternalStore, also as the server snapshot (an untouched id readsabsenton both workers):const status = useSyncExternalStore( (onChange) => FuzzySearchManager.subscribeWorkerStatus(workerId, onChange), () => FuzzySearchManager.getWorkerStatus(workerId), () => FuzzySearchManager.getWorkerStatus(workerId), )Whether a manager would get its workers is the engine's decision, asked without side effects.
FuzzySearchManager.workerAcquisitionOf(workerId, workerUrls?)answers exactly when a manager of that worker id with thoseworkerUrlsgets both workers — the rule the acquisition itself follows, so a host never re-derives it from the status. Per worker (the index worker first, the Levenshtein worker only when the index worker can be had): the id holds it live (loadingorready, whatever its URL), or Web Workers exist and the resolved URL is notnulland not remembered failed. It returns a frozenWorkerAcquisition:const acquisition = FuzzySearchManager.workerAcquisitionOf('products', { indexWorker: '/assets/indexWorker.mjs' }) acquisition.allowed // whether such a manager gets both workers acquisition.urls // the resolved URLs (ResolvedWorkerUrls, frozen) acquisition.unresolved // the workers whose URL cannot be resolved here ('index' first; 'levenshtein' only when the index worker can be had) acquisition.workers // per worker, whether such a manager gets it: { index, levenshtein } (allowed is both)workersanswers per worker by the same decision:index— the id holds it live, or it may be created;levenshtein— a live one always, a new one only when the index worker can be had.allowedis exactlyworkers.index && workers.levenshtein. Every call needs the index worker, so a host may hold a manager wheneverworkers.indexistrue: withworkers.levenshteinfalse,searchWithStages()(below) answers with the index stage and reports the Levenshtein stageunavailable, whilesearch()rejects as it always has.The query is pure: it records nothing in the status, notifies no one, logs nothing, creates and fetches nothing, and parses no URL for a repeated input. The answer stays the same object while its values are unchanged, so it (or its
allowed) is auseSyncExternalStoresnapshot withsubscribeWorkerStatus; a host that wants to tell the user about an unresolvable URL reportsunresolveditself:const acquisition = useSyncExternalStore( (onChange) => FuzzySearchManager.subscribeWorkerStatus(workerId, onChange), () => FuzzySearchManager.workerAcquisitionOf(workerId, workerUrls), () => null, )FuzzySearchManager.canAcquireWorkers(workerId, workerUrls?)is deprecated (removed in the next major version): it returns the sameallowedbut also records eachunresolvedworker in the id's status as thecreatefailure withurl: nullthat an acquisition records (subscribers hear it in a microtask).States.
absent→loadingwhen a manager creates the worker;loading→readyat the worker's first message (the handshake every manager sends);failedwith phasecreate(theWorkerconstructor threw, or no URL could be resolved),load(the worker's ownerrorevent fired before its first message, and no worker from that URL has answered under the id — the script could not be fetched or evaluated) orrun(its ownerrorevent fired after it had answered, or before the first message of a worker re-created from a URL that has answered under the id: a URL that has answered once is proven and never becomes a remembered load failure).A failed worker is terminated at once. The engine listens to every worker it creates (a failure is heard even when no manager is alive), terminates the failed worker, drops it, clears the shared index and dataset cache of the id, and settles every wait on it at once instead of letting it wait out its timeout: a
loadfailure rejects them with aWorkerFailureError(worker,workerId,workerUrl,phase, andrefusal:nullfor a death,'dataset'or'breaker'for a refusal described below), arunfailure hands a waiting search to the re-send below. This includes the waits of a manager disposed meanwhile, so a search that joined its index build never waits out the build timeout.A URL that failed to create or load is never fetched again for the page's life under that worker id and worker: managers whose URL is remembered get no such worker (their
search()rejects as when the workers are unavailable), so a broken URL costs one fetch, not one per keystroke. Any other URL is a fresh attempt — the first manager (or request) with a different URL creates the worker from it.A search survives a worker that stops after answering. When a worker dies with phase
runwhile asearch()waits on it, the engine re-creates the worker from that search's manager's URL and re-sends the search's part: the index part rebuilds the index on the new worker first, the Levenshtein part re-sends the dataset. A worker handles its requests one at a time in the order they were posted, so the engine blames the death on the dataset of the oldest request the worker had not answered (the one it was processing; an idle death blames none). A part of that dataset spends its one re-send per worker kind and search, so a search during which both workers die still resolves (each gets its own re-send); a part of any other dataset waiting on the same worker — another box's search under a shared worker id, or the previous options' search of the same box — is re-sent without spending it, so a dataset that never crashed a worker is never recorded below. Searches waiting together share the re-created worker and its rebuild. Every engine user gets this policy — directsearch()callers and theuseFuzzySearchhook alike (the hook does not turn arun-phaseindexWorker:errorinto itserror; a final failure reaches it through the rejected search).A dataset a worker stops on twice is not sent to it again. When the same worker dies with phase
runagain during the re-send, the search rejects with aWorkerFailureError(phaserun) whose message ends with(it stopped again when the request was re-sent to a new worker, so the engine sends this dataset to the <Index | Levenshtein> Worker no more; a changed dataset is sent once the worker's circuit breaker lets requests through), and the dataset is published instatus.<worker>.crashedDatasets({ dataHash, failure }, the list never shrinks). Any latersearch()orprebuildIndex()of that dataset that needs that worker rejects at once with the same message, posting nothing and creating no worker (the index worker is needed by every search, the Levenshtein worker whenenableLevenshtein), so a deterministic crash (an out-of-memory index build, say) costs two crashes per dataset, not two per keystroke; a changed dataset is sent once the worker's circuit breaker (below) lets requests through, and a search withenableLevenshtein: falsestill gets the index stage's answer after a Levenshtein crash. The error'srefusalis'dataset'(on the final rejection too).rebuildIndex(), the explicit forced rebuild, is never refused.search()resolves with the full answer of every enabled stage or rejects; it never resolves with one stage's rows only.searchWithStages()(below) is the same search that answers with the index stage when only the Levenshtein worker cannot serve, and says so.A circuit breaker per worker bounds content that keeps crashing it. A crash caused by content every new dataset keeps (a pathological row in an append stream, a size limit) would otherwise cost two crashes, two re-created workers and two whole-dataset posts per dataset change. So the final crash that records a dataset also opens that worker's breaker under the worker id (
status.<worker>.breaker, frozen):status.levenshtein.breaker // { open: true, probing: false, openedAt: 1767225600000, nextProbeAt: 1767225605000, // failure: { phase: 'run', url, message }, openings: 1 }While
openistrue, everysearch()needing that worker and everyprebuildIndex()(index worker) rejects at once — before any worker is created or anything is posted — with aWorkerFailureError(phaserun,refusal'breaker') whose message ends with(the engine's circuit breaker for the <Index | Levenshtein> Worker of this workerId opened after this failure (opening <n>): it refuses every request until <ISO time>, then lets one request through to probe it). The window is 5 s for the first opening and doubles with every consecutive opening (10 s, 20 s, 40 s, 80 s, 160 s), capped at 300 s;nextProbeAtis when it ends (epoch milliseconds, likeopenedAt). AtnextProbeAtthe engine publishesopen: false, and the next call needing that worker is let through as the probe (openandprobingaretruewhile it is in flight, and other calls are refused with… it refuses every request while the one request it let through to probe it is answered)).probingistrueexactly while the probe is in flight, soopen && !probingis the window itself, and a host that readsprobingwaits for the probe's outcome — the next status publication — instead of treating its own pending call as refused. When the probe resolves the breaker closes (breakerbecomesnull, and the next opening starts from 5 s again); when the probe's own request kills the worker (a search after its re-send, aprebuildIndex()at its one death) the breaker opens again with the next window; when the probe ends any other way the next call is the probe. So such content costs at most one opening (two crashes) per window, however often the dataset changes. Only the workers a call needs are checked: with only the Levenshtein breaker open, a search withenableLevenshtein: falsegets the index stage's answer. Calls already running when the breaker opens finish their parts;rebuildIndex()is never refused and does not touch the breaker.crashedDatasetskeeps listing only the datasets whose re-send died. AprebuildIndex()death while the breaker is closed opens nothing. A refused call reads its dataset only when a dataset with the same item count and search fields is recorded for that worker (equal datasets have equal shapes), so the calls of an append stream refused during a window hash nothing; a refusal ofprebuildIndex()is logged at debug level (a refusedsearch()is not logged; its caller gets the rejection).A detailed search reports what ran.
searchWithStages(query, items, searchFields, options?)issearch()with the Levenshtein worker optional — one implementation, so it runs, re-sends and rejects exactly assearch()does except when only the Levenshtein worker cannot serve the call. The engine's admission decides per worker: the index worker is required (its refusal or failure rejects assearch()does), and when the Levenshtein worker's breaker refuses the dataset, the dataset is recorded for it, it cannot be had, its script fails to load, or the call's re-send dies on it again, the call resolves with the index stage's answer:const { results, stages } = await manager.searchWithStages('aple', items, ['name']) stages.index // { ran: true } stages.levenshtein // { ran: true } or { ran: false, reason: 'breaker' | 'dataset' | 'crashed' | 'unavailable' | 'disabled' | … }Reasons (frozen reports):
blank(blank query; both stages),disabled(the stage is off in the options),disposed(the manager was disposed first; results[]),superseded/contended(index stage: a newer dataset of this manager overtook the search, or the shared index still held another dataset after one rebuild),breaker(the Levenshtein breaker refused the dataset at admission — its window, or another call is its probe),dataset(recorded for the Levenshtein worker before the call),crashed(recorded during the call) andunavailable(the Levenshtein worker could not be had, or its script failed to load).blank,disabled,dataset,crashedandunavailableare permanent for the manager, the dataset and the options;breaker,supersededandcontendedare temporary — an answer carrying one of them is not the dataset's complete answer and should be searched again later. A Levenshtein failure outside that list (its 30 s timeout, an error reply) rejects as insearch(). A call that resolves closes the breaker probe of each worker whose stage served it and ends the probe of a Levenshtein worker whose stage could not serve (that breaker re-opens on a death blamed on the call's dataset, otherwise the next call probes); it emits the samesearch:*events and updates the same stats.Outside a search, a worker that failed after answering is recreated by the next request that uses it. Nothing is remembered for a
runfailure: the nextsearch()/prebuildIndex()/rebuildIndex()of any manager of the id that uses the worker and is not refused creates it again from that manager's resolved URL (and rebuilds the index on it:prebuildIndex()skips only when the id's shared index already holds the data).prebuildIndex()andrebuildIndex()re-send nothing: a worker death during them rejects them. Every request first re-acquires, of the workers it uses, the ones its manager lacks, not only the constructor: a search uses the index worker and, withenableLevenshtein, the Levenshtein worker;prebuildIndex()andrebuildIndex()use the index worker alone; a search part's re-send re-acquires only the worker that died. A worker a request does not use is never created for it, so a search withenableLevenshtein: falseneither re-creates a stopped Levenshtein worker nor needs one. The constructor prepares both workers.Events stay as they were, with more in them.
workers:initialized(once per manager, at the end of its constructor's setup) keepsindexWorker/levenshteinWorkerand addsworkerId,workerUrls(the manager's resolved URLs) andstatus. For a worker's ownerrorevent,indexWorker:error/levenshteinWorker:errorstill deliver the originalEvent, now carryingworker,workerId,workerUrlandphase; an error reply to one request with no waiter is still the{ requestId, error, processingTime }record.
Text normalization
The engine compares raw text unless you pass normalize. With it, both workers index normalizeText(String(item[field] ?? ""), normalize) for every search field and read every query as normalizeText(query, normalize); all matching (candidates, Jaro-Winkler, containment, Levenshtein, the "and" token check) then runs on those texts. Results still carry the original items, originalIndex and matchedFields (names of the original fields).
import { FuzzySearchManager } from "@aiquants/fuzzy-search"
const manager = new FuzzySearchManager<{ label: string }>({ normalize: { convertRomaji: true } })
const items = [{ label: "ヤマダ" }, { label: "やまだ" }, { label: "ヤマダ" }, { label: "山田" }]
const results = await manager.search("yamada", items, ["label"])
// ヤマダ, やまだ and ヤマダ are found (each normalizes to ヤマダ); 山田 is not (kanji have no reading here)- Absent (
undefined): no normalization — the engine behaves exactly as 2.2.x (same matching, same index keys). - An object: omitted flags take
DEFAULT_NORMALIZE_OPTIONS(convertCase,convertWidth,convertKana,normalizeVariantson;convertRomaji,ignoreSpaces,ignoreSymbols,ignoreLongVowelsoff), so{}already makes hiragana ↔ katakana and half-width ↔ full-width equal. It is accepted by the constructor, by the per-calloptionsofsearch()/prebuildIndex()/rebuildIndex()(a call's ownnormalizekey wins,normalize: undefinedturns it off for that call) and byuseFuzzySearch. - Validated on the main thread by
resolveNormalizeOptions(value, "[fuzzy-search] normalize"): the constructor throws, andsearch()(even for a blank query) /prebuildIndex()/rebuildIndex()reject, with a RangeError before any worker is touched —[fuzzy-search] normalize must be a plain object; got 5,[fuzzy-search] normalize keys must be one of "convertCase", "convertWidth", "convertKana", "convertRomaji", "normalizeVariants", "ignoreSpaces", "ignoreSymbols", "ignoreLongVowels"; got "convertCaps",[fuzzy-search] normalize.convertCase must be a boolean; got 1,[fuzzy-search] normalize.convertKana must be true when convertRomaji is true; got false. The resolved options travel to the workers inside the request options (structured clone). - A query that normalizes to nothing (
"!?"underignoreSymbols) returns no results. - Shared workers: the dataset key includes the resolved flags, so managers of one
workerIdwith different (or no) normalization never reuse each other's index. - Case:
normalize.convertCasefolds case first;caseSensitivethen applies to the normalized texts. - Positions: the engine reports no character positions, only
matchedFields(names of the original fields). The position-mapped forms that map a range of a normalized form back to the source text (toMappedFormsInto,MappedText) are part of the normalize entry's query compilation forms.
Equivalences (each line run through the real normalizeText):
| Input | Options | Normalized |
| --- | --- | --- |
| yamada, YAMADA, やまだ, ヤマダ, ヤマダ | { convertRomaji: true } | ヤマダ |
| Tōkyō | { convertRomaji: true } | トーキョー |
| toukyou, トウキョウ | { convertRomaji: true } | トウキョウ |
| ko-nsuta-chi, コーンスターチ | { convertRomaji: true } | コーンスターチ |
| Matuyama | { convertRomaji: true } | マツヤマ |
| ヤマダ, ヤマダ | {} | ヤマダ |
| ABC | {} | abc |
| Straße | {} | strasse |
| ポテト-25kg, ぽてとー25kg | {} | ポテトー25kg |
| ㄱㅏ, 가 | {} | 가 |
| コーン, コオン | { ignoreLongVowels: true } | コン |
| Shinʼya | { ignoreSymbols: true } | shinya |
| 山田 太郎 | { ignoreSpaces: true } | 山田太郎 |
Limits:
- Kanji readings are out of scope:
山田stays山田. To find kanji by reading, put the reading in the data (a reading field such as{ label: "山田", reading: "やまだ" }) and search both fields. - One reading per text: fields and queries are read alike, with the option-text reading of romaji (
ti/tu/di/du/woasチ/ツ/ヂ/ヅ/ヲ, sotisshureadsチッシュ). The alternative query readings a component may compile (ティforti) are not applied by the engine. - A normalized text is a whole-text key: normalization reads characters from their neighbours, so a fragment can read differently inside the word (under
convertRomaji,ationreadsアチオンwhilenationreadsナチオン;abcreadsアbc). Such fragments still meet through the fuzzy stages, but not as literal substrings.
Cost: the normalizer runs inside each worker, at index build (every field value once) and once per query. The worker scripts do not carry it: each worker loads worker/normalizer.mjs (25.4 KB, 10.1 KB gzip) once, at the first request with the option, from the URL the manager resolves (see Worker setup); a request whose normalizer cannot be loaded rejects with [fuzzy-search] the normalizer could not be loaded from <url>: …, and search() /
prebuildIndex() / rebuildIndex() with the option reject before touching a worker when no normalizer URL can be resolved ([fuzzy-search] normalize: no URL could be resolved for the normalizer script (specify options.workerUrls.normalizer explicitly)). The main entry validates the option with the ./normalize entry's options module, which it shares in the ESM build.
The normalize entry (@aiquants/fuzzy-search/normalize)
The same normalizer, importable on its own: no React and no worker in its import graph, ESM split per module so a bundler keeps only what is used (measured with esbuild, minified: hiraganaToKatakana alone 2.6 KB, resolveNormalizeOptions 1.4 KB, romajiToKatakana 8.5 KB, normalizeText 25.4 KB, getNormalizedVariants 30.6 KB, the whole entry 43.7 KB).
import { getNormalizedVariants, hiraganaToKatakana, katakanaToHiragana, normalizeText, romajiToKatakana } from "@aiquants/fuzzy-search/normalize"
normalizeText("ヤマダ") // "ヤマダ"
normalizeText("yamada", { convertRomaji: true }) // "ヤマダ"
romajiToKatakana("yamada") // "ヤマダ"
hiraganaToKatakana("やまだ") // "ヤマダ"
katakanaToHiragana("ヤマダ") // "やまだ"
getNormalizedVariants("ヤマダ") // ["ヤマダ", "ヤマダ", "やまだ"]Documented API: normalizeText(text, options?), getNormalizedVariants(text, options?), normalizeJapaneseVariants(text), hiraganaToKatakana(text),
katakanaToHiragana(text), romajiToKatakana(text), resolveNormalizeOptions(options, name) (name is the setting as the caller names it, owner first,
and starts every message — the engine passes [fuzzy-search] normalize, a library validating its own setting passes e.g. [my-lib] normalizeOptions),
DEFAULT_NORMALIZE_OPTIONS, and the types NormalizeOptions / ResolvedNormalizeOptions. Every function rejects a non-string text with a RangeError, and
normalizeText / getNormalizedVariants / resolveNormalizeOptions reject invalid options the same way ([fuzzy-search] normalizeText: text must be a string; got 42,
[fuzzy-search] normalizeText: options.convertCase must be a boolean; got 1).
Under moduleResolution: node10 (which ignores the exports map) the normalize and react subpaths resolve their declarations through the package's typesVersions.
Query compilation forms
The entry also exports the forms a host compiles its own queries over — documented API, covered by semver like the functions above (a change to any of them is a major release):
| Group | Names |
| --- | --- |
| Comparison forms of a text | FORM_TIERS, toTextForms, toTextFormsInto, toMappedFormsInto, InteriorPairs, interiorEndAt, distinctForms, trimmedRange; types FormSet, FormKind, FormTier, TextForms, MappedForms, MappedText, PartialReadings, InteriorCursor, DistinctForm |
| Query reading and needles | readQuery, segmentQuery, isBlankQuery, queryNeedles, occurrenceTier, isAnswerTier; types QuerySegmentation, QueryNeedle, Tier |
| Per-engine readings | ENGINE_READING_CONVENTIONS, engineForms, engineFieldsOf, engineWordAnchorsOf, placeBesideOthers, perEngineReading, LONE_WORD_PLACE; types EngineForms, EngineReadingConvention, EngineReadings, EngineWordAnchors, QueryWordPlace |
| Grapheme clusters | isClusterBoundary, clusterCountOf, holdsMultiUnitCluster |
Each name's contract — its inputs and outputs, what each tier means, the order FORM_TIERS guarantees, what typing changes, the
reading-interior protocol of InteriorPairs, the shapes of the Hangul IME stand-ins and the status of the jamo members — is in
Query compilation forms, shipped in the package as docs/query-compilation-forms.md.
While a query is being typed (queryNeedles(query, options, true)), the needles also hold tier-2 Hangul IME stand-ins:
what a Korean two-set IME can still turn the syllable it is composing into, each an answer-tier reading (literal,
canonical, the long-vowel fold and each romaji reading) of a real continuation of the query, lone compound consonants
included (ㄱ also stands for ㄳ, ㄳ for ㄱ + ᄉ), so a separator ignoreSpaces / ignoreSymbols removes never
lets a compound compose onto the syllable it committed (가 ㄱ stands in as 가 ᆪ, not 갃). They are jamo texts,
compared with the jamo image of every form a host stores; the page gives their shapes and how to match them in
composed forms. FormSet.literalJamo / canonicalJamo (the images of the literal and canonical forms) are own data
properties like every other member. distinctForms(forms, textOf, tiers) reads only the kinds in tiers, which must
be FORM_TIERS or a subsequence of it (else a RangeError naming the first offending entry).
Nothing else is exported: the entry exports 51 names (the 10 documented normalization names, and the 24 values and 17 types of this table), and its export list is pinned by a test, so it cannot grow or shrink silently.
Configuration Options
Top-level options (FuzzySearchOptions, all optional in the constructor / hook):
| Option | Default | Description |
| --- | --- | --- |
| threshold | 0.4 (Index) / 0.3 (Levenshtein) | Minimum similarity score (0-1). Propagated to worker options if not explicitly overridden |
| caseSensitive | false | Case-sensitive matching |
| multiTermOperator | "or" | How whitespace-separated query tokens combine. "and" requires every token as an exact substring |
| ngramSize | 2 | N-gram size for the index. Shorter queries fall back to substring matching over indexed keys |
| minNgramOverlap | 1 | Per-token lower bound of the n-gram overlap threshold in the accurate candidate strategy |
| sortBy | "relevance" | "score" | "relevance" (score + bonuses for fields meeting the threshold and actual partial matches) | "original" (input order) |
| sortOrder | "desc" | Sort direction. Not applied when sortBy is "original" (always input order) |
| enableIndexFiltering | true | Run the index (candidate) stage |
| enableLevenshtein | true | Run the Levenshtein stage (full scan) |
| parallelSearchStrategy | "balanced" | Result merging: "balanced" (merge by score) | "index-first" | "levenshtein-first" |
| customWeights | {} | Per-field score multipliers, e.g. { name: 2.0 } |
| debounceMs | 300 | Debounce for the React hook |
| autoSearchOnIndexRebuild | true | Re-run the current query after the hook rebuilds the index |
| workerId | "default" | Worker sharing/isolation key |
| workerUrls | – | Explicit script URLs { indexWorker, levenshteinWorker, normalizer } (required for CJS; normalizer only with normalize); resolved by the worker URL rule |
| logLevel / logger | WARN / console | Logging control. The logger object stays on the main thread (it is not sent to workers) |
| normalize | – | Text normalization of every field value and query in both workers (see Text normalization); absent = raw text, as 2.2.x |
Worker tuning options: indexWorkerOptions (strategy: "fast" | "accurate" | "hybrid", threshold, ngramOverlapThreshold, minCandidatesRatio, maxCandidatesRatio, jaroWinklerPrefix, maxResults, relevanceFieldWeight, relevancePerfectMatchBonus, containmentScoreBase, containmentCoverageWeight) and levenshteinWorkerOptions (threshold, lengthSimilarityThreshold, lengthDiffPenalty, partialMatchBonus, maxResults, relevanceFieldWeight). Explicit 0 values are honored.
The index stage scores each field as max(JaroWinkler, containment): when the case-normalized, trimmed query is a literal substring of the field value, the containment score containmentScoreBase + containmentCoverageWeight * (trimmedQueryLength / valueLength) (defaults 0.7 / 0.3; an identical string scores exactly 1.0) rescues matches the Jaro-Winkler window cannot reach.
The scoring query is trimmed once up front, so Jaro-Winkler, containment, and the perfect-match equality check are all invariant to surrounding whitespace. In sortBy: "relevance" mode, the perfect-match bonus (relevancePerfectMatchBonus) is granted only on explicit string equality — not inferred from a high score.
For queries with more than one token, each field is scored against the whole query and against each token individually, and the best of those wins.
Without this, tokens spread across different fields (e.g. "佐藤 商事" where one field holds "佐藤 商" and another "商事ファシリティーズ") leave every field failing to contain the whole query, so the containment bonus never applies and the whole query's length dilutes the similarity.
The relevance bonus follows the same rule: a field matching any single token qualifies — by string equality in the index stage (relevancePerfectMatchBonus), by substring containment in the Levenshtein stage (partialMatchBonus).
Candidate retrieval treats whitespace-separated query tokens with OR semantics at every stage: the fast strategy collects n-grams from the whole query and from each token individually, the candidate-overflow filter and the accurate strategy evaluate n-gram overlap per token (an item survives if any one token clears its threshold), and tokens shorter than ngramSize — which cannot produce n-grams — fall back to substring matching over indexed keys instead of being silently dropped.
A mixed query like "A 粉末" therefore keeps the items that only match the short token, and "佐藤 郎" returns the union of both tokens' hits rather than dropping the single-character token's side entirely (all three strategies return the same union).
API Reference
Main entry (@aiquants/fuzzy-search)
FuzzySearchManager<T>—search(),searchWithStages(),prebuildIndex(),rebuildIndex(),getStats(),getCurrentStats(),resetStats(),getSearchHistory(),updateSearchHistory(),on()/off()(search lifecycle events),dispose()- Statics:
FuzzySearchManager.resolveWorkerUrls(workerUrls?),FuzzySearchManager.workerAcquisitionOf(workerId, workerUrls?),FuzzySearchManager.canAcquireWorkers(workerId, workerUrls?)(deprecated),FuzzySearchManager.getWorkerStatus(workerId?),FuzzySearchManager.subscribeWorkerStatus(workerId, listener),FuzzySearchManager.defaultWorkerUrls WorkerFailureError,DEFAULT_FUZZY_SEARCH_OPTIONS,isSearchResultItem,LogLevel- Types:
FuzzySearchOptions,PartialFuzzySearchOptions,SearchResultItem,SearchResults,StagedSearchResults,SearchStages,SearchStageReport,SearchStageSkipReason,SearchStats,UseFuzzySearchOptions,ILogger,WorkerStatus,WorkerSlotStatus,WorkerState,WorkerFailure,WorkerFailurePhase,CrashedDataset,WorkerBreaker,WorkerRefusal,WorkerUrls,ResolvedWorkerUrls,WorkerAcquisition,EngineWorkerKind,WorkerErrorEvent,WorkerErrorDetails,WorkersInitializedEvent, …
React entry (@aiquants/fuzzy-search/react)
useFuzzySearch(items, searchFields, options?)— main search hookuseSearchHistory(maxHistory?)— standalone search-term history
Normalize entry (@aiquants/fuzzy-search/normalize)
normalizeText,getNormalizedVariants,normalizeJapaneseVariants,hiraganaToKatakana,katakanaToHiragana,romajiToKatakana,resolveNormalizeOptions,DEFAULT_NORMALIZE_OPTIONS- Types:
NormalizeOptions,ResolvedNormalizeOptions - The query compilation forms (see Query compilation forms)
Worker entries
@aiquants/fuzzy-search/worker/indexWorker@aiquants/fuzzy-search/worker/levenshteinWorker@aiquants/fuzzy-search/worker/normalizer(loaded by both workers for thenormalizeoption)
Performance Tips
- Pre-build the index with
prebuildIndex()(the React hook does this automatically whenitemschanges) - Keep the
itemsarray reference stable between searches — the dataset content hash is memoized per array reference, so a repeated call with the same array hashes nothing (mutating that array in place is not detected: pass a new array or callrebuildIndex()) - Use separate
workerIds for components searching different datasets - Set
maxResultsin the worker options for very large result sets - Adjust
threshold— lower values return more (and noisier) results
Browser Support
- Chrome / Chromium 60+
- Firefox 55+
- Safari 11+
- Edge 79+
License
MIT License
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
