@promptev/context-engine
v1.0.1
Published
Promptev Context Engine: language-agnostic ingestion + hybrid retrieval (TypeScript)
Maintainers
Readme
Built by Promptev. Apache-2.0. The Python twin,
promptev-context-engine, shares the same Postgres schema, so a corpus ingested by either is searchable by the other.Self-host it, or try it in the cloud without hosting anything: bring your own database and model keys and explore promptev cloud.
@promptev/context-engine is a TypeScript library for applications that
answer questions over documents and act through tools, and that have to say
who may see what. It ingests files into the Postgres you already run,
retrieves with full-text, trigram and vector legs fused together (a graph leg
optionally), and puts two things where other stacks leave them to application
code: the access-control list goes into the SQL predicate every retrieval leg
shares, so a document a caller may not see is never fetched, never ranked and
never reaches a reranker; and PII redaction runs before a hit is sent to a
reranker, before a tool result is written to its audit row, and (if you ask)
before a chunk is embedded.
The models are yours. Hand the engine one function and every model call it makes (extraction, vision transcription, graph building, map-reduce, compute) goes through your own harness, with your keys, retries, limits and cost accounting. A small OpenAI-compatible client is built in for a quick start.
It is a library, not a server: no UI, no orchestration, no opinion about your
framework. It never calls a service you did not configure and implements no
auth, billing or credential storage of its own; it exposes hooks (onUsage,
onError, onProgress) and injection points (auth, principals, scope)
for your application to wire up. It does not prevent hallucination or
guarantee grounding, and its compute() action (model-written code) is off by
default.
Full documentation: https://promptev.ai/documentation/context-engine/ · in this repo: access control · the ACL-recall benchmark · cookbook · model adapters
Install
npm install @promptev/context-engine pgNeeds Node.js 22+ and Postgres with the vector (pgvector 0.8+),
pg_trgm and unaccent extensions, created by context-engine migrate when
the role may CREATE EXTENSION. pg is a peer dependency, and it, like
everything else here, is MIT/BSD/Apache, so a bare install pulls no copyleft.
No provider SDK is installed: the built-in client is plain HTTP, and a host
model brings its own.
Optional peers, install only what you use:
| Peer | Needed for |
|---|---|
| pg | the database driver; you almost certainly want this |
| hono / express / fastify | createHonoRouter / createExpressRouter / createFastifyPlugin (Fastify uploads also need @fastify/multipart on your app) |
| graphology + graphology-communities-louvain | mode: "graph" (community detection) |
| neo4j-driver | mode: "graph" with a neo4jUri; optional, see Graph mode |
| sharp, @napi-rs/canvas | page rasterization for scanned pages and images, transcribed by whichever model you bring |
| tesseract.js | OCR fallback |
| @modelcontextprotocol/client | calling third-party MCP servers as mcp tools; see Connecting MCP servers |
| @modelcontextprotocol/server + @modelcontextprotocol/node | createMcpApp / context-engine mcp |
| danfojs-node + isolated-vm | engine.compute(): dataframes and the sandbox |
| pdfjs-dist | structure-preserving PDF text (tables stay tables) |
Upgrading an existing deployment: MIGRATING.md.
Quick start
Migrate the database first. The engine needs a Postgres with pgvector and its own tables. To try it locally, start one and migrate it:
docker run -d --name context-engine-pg -p 127.0.0.1:5432:5432 \
-e POSTGRES_USER=user -e POSTGRES_PASSWORD=pass -e POSTGRES_DB=mydb \
pgvector/pgvector:pg16
until docker exec context-engine-pg pg_isready -h localhost -U user -d mydb; do sleep 1; done
npx context-engine migrate --database-url postgresql://user:pass@localhost:5432/mydb --dim 1536The second line waits until Postgres accepts connections: docker run -d
returns before it does. The port is published on loopback only. If port
5432 is already taken, publish another one (-p 127.0.0.1:5433:5432) and
use it in the URL (localhost:5433).
migrate creates the extensions and the tables, and is safe to run again
(see Provision the database). On a database that
was never migrated, the first call throws DatabaseNotMigrated, and its
message is the command to run:
DatabaseNotMigrated: This database has no Context Engine tables, or this role cannot see them. Run the migration once, then try again: npx context-engine migrate --database-url <your database URL> --dim 1536DatabaseNotMigrated is an Error; the HTTP routers answer it with 503 and
the MCP tools with one fixed sentence.
The example below uses top-level await, so the project needs
"type": "module" in its package.json.
Bring your own model (recommended)
Run your own model stack (keys, routing, rate limits, tracing, cost) and hand
the engine three functions. Every model call, every embedding request and
every rerank then goes through them. This is complete and runs as written;
the three stubs are what you replace with calls on your own provider (or copy
a ready adapter from examples/model_adapters/):
import {
ContextEngine,
ContextEngineConfig,
type EmbedKind,
type ModelReply,
type ModelRequest,
} from "@promptev/context-engine";
async function myModel(request: ModelRequest): Promise<ModelReply> {
// request.purpose is one of structured_extraction, entity_extraction,
// community_summaries, vision_pages, vision_image, map_reduce, compute.
// request.system / request.user are the prompt; request.images are PNG
// bytes for the vision purposes; request.responseSchema, when set, is a
// strict JSON Schema you can hand to a provider that enforces one; and
// request.maxTokens / temperature / thinkingBudget are the engine's asks.
// Stub: "{}" does not satisfy request.responseSchema, so the structured
// purposes fail until this returns your provider's real reply.
const text = request.responseSchema ? "{}" : "";
return {
text,
tokens: { input: 0, output: 0 }, // what your provider billed
mode: "prompt", // "native" only when it enforced request.responseSchema
};
}
async function myEmbedder(texts: string[], { kind }: { kind: EmbedKind }): Promise<[number[][], number]> {
// kind is "document" or "query". One vector per text, config.embedding.dim
// wide, and the tokens billed. Stub: a unit vector, never all zeros
// (cosine distance to a zero vector is undefined and breaks search).
return [texts.map(() => [1, ...new Array<number>(1535).fill(0)]), 0];
}
async function myReranker(query: string, docs: string[]): Promise<number[]> {
return docs.map((_, i) => i); // indices into docs, most relevant first
}
const config = new ContextEngineConfig({
databaseUrl: "postgresql://user:pass@localhost:5432/mydb",
// Required with a host embedder too: the identity the database records
// (model name, dimension) and the batch limits. No apiKey needed.
embedding: { provider: "openai", model: "text-embedding-3-small", dim: 1536 },
});
const engine = new ContextEngine(config, { model: myModel, embedder: myEmbedder, reranker: myReranker });
const report = await engine.ingest({
text: "Annual leave accrues at two days per month.",
name: "leave-policy",
sourceId: "hr-handbook",
acl: ["group:hr"],
});
console.log(report.documents[0].status); // "completed"
const result = await engine.search("how much annual leave", { principals: ["group:hr"] });
console.log(result.hits.map((hit) => hit.chunkText));
await engine.aclose();The engine keeps the prompts, the JSON Schemas, reply validation with one
repair retry, redaction before anything leaves, and metering. mode: "native"
promises your provider enforced the schema; leave it at the default "prompt"
when it did not, and the engine validates and repairs. With a model,
config.llm, config.visionLlm and graph.extractionLlm are optional, and
vision is on. A model that cannot read images should throw for vision_*;
those pages end as vision_failed. A refusal goes in refusal, never in
text, so it is never stored as a document's data.
embedder is a function (texts, { kind }) resolving to [vectors, tokens],
or any object with an embed method of that shape; the dimension check runs
on every reply. reranker is (query, docs) resolving to number[]; it runs
on the widened candidate window (config.reranker.candidates, 50), and the
engine drops out-of-range and duplicate indices and appends any it omitted,
so it can reorder hits and never drop them. Each of the three may return a
value or a promise.
Your harness owns the call. The engine builds each request (prompt,
schema, images, redaction) and checks the reply. Everything about making the
call is yours: retries, timeouts, rate limits, concurrency, caching, cost. The
engine does not retry a host call (it repairs a reply that breaks the schema
once, and re-asks for pages a vision reply omitted) and does not time it out;
whatever the host throws is final for that call site. So you can size your
limits: the engine sends up to 8 vision calls per document, 5 map_reduce
calls by default (20 at most) and 4 embedding batches at once, and concurrent
ingests add up. Which model writes compute code, most specific first: a
per-call model on compute() / computeOverFrames(), then a per-call
modelCfg, then the engine's model, then config.llm. A call you hand to
a provider batch API can throw ModelDeferred instead of blocking; see
Batch mode.
Or use the built-in default
Pass no model or embedder and the engine calls the endpoints named by
config.llm and config.embedding itself, through one OpenAI-compatible
client over plain HTTP (POST {baseUrl}/chat/completions and
POST {baseUrl}/embeddings):
import { ContextEngine, ContextEngineConfig } from "@promptev/context-engine";
const config = new ContextEngineConfig({
databaseUrl: "postgresql://user:pass@localhost:5432/mydb",
embedding: { provider: "openai", model: "text-embedding-3-small", apiKey: "sk-..." },
// The model for structured extraction, compute, map_reduce and query_meta.
// Optional, and unused when the engine is given a host model:
llm: { provider: "openai", model: "gpt-4o-mini", apiKey: "sk-..." },
// Optional: the model that transcribes scanned pages and images.
visionLlm: { provider: "openai", model: "gpt-4o-mini", apiKey: "sk-..." },
});
const engine = new ContextEngine(config);provider is one of three:
| provider | endpoint | auth |
|---|---|---|
| openai | https://api.openai.com/v1 unless baseUrl says otherwise | Authorization: Bearer |
| azure_openai | baseUrl required: the Azure v1 path, https://<resource>.openai.azure.com/openai/v1/ | api-key header |
| custom | baseUrl required: any OpenAI-compatible root | Authorization: Bearer when apiKey is set |
provider names the wire format, not the vendor: custom with the endpoint
root as baseUrl reaches Ollama, vLLM, Groq, Mistral, a LiteLLM proxy and
anything else that speaks the OpenAI chat and embeddings format. The
per-vendor table is under Use any model provider.
The client retries 408, 409, 429 and 5xx with full-jitter backoff up to
maxRetries (2), honouring Retry-After; any other 4xx is final. On a
search query's embedding the wait is capped at
search.queryEmbedRetryAfterCapS (5 seconds) with
search.queryEmbedMaxRetries (1), so a rate-limited key cannot hold a search
for minutes. A successful reply that carries no choice, or an embeddings reply
that does not hold exactly one vector per input with each index used once,
throws OpenAICompatReplyError rather than passing as an empty answer. It
never
follows a redirect: a 3xx from a provider endpoint would carry the key and the
prompt wherever Location points, so it fails the call with that status. It
claims an enforced (native) JSON schema only on OpenAI proper and Azure;
every other endpoint is reported as prompt and validated, because accepting
response_format is not enforcing it. Set schemaEnforced: true on an llm
config for an endpoint that does enforce it (vLLM with guided decoding, say),
or false to validate everywhere. thinkingBudget is ignored by this client,
since thinking control is provider-specific; a host model sees it.
Both paths can run on a fetch you supply instead of the engine's own:
new ContextEngine(config, { fetchFactory, mcpLookup }). fetchFactory(purpose)
returns a fetch for "llm" | "embeddings" | "reranker" | "tool_http" |
"mcp_oauth", or null to keep the engine's; it must be long-lived and
host-owned, since the engine never closes anything behind it. mcpLookup is
the resolver MCP connections open through. A host model, embedder or
reranker never calls it.
Use any model provider
The engine works with any model provider. There are two ways in, and one rule:
model/embedder/reranker(recommended): any SDK, any vendor, any local model, your own gateway. You keep the provider's native schema enforcement, prompt caching, thinking control, and your own retries, routing and tracing. The engine never sees a key.- The built-in OpenAI-compatible client:
provider: "openai"and"azure_openai"only fill in the URL and the auth header;provider: "custom"plusbaseUrlreaches anything that speaks the OpenAI chat/embeddings format.
The rule: provider names the wire format, not the vendor. provider: "anthropic",
"gemini", "vertex_ai" or "bedrock" is refused at config time, on purpose,
with a message pointing at the two ways above; a vendor URL table inside the
engine would go stale, and model gives you the vendor's real SDK instead.
| Vendor or runtime | Recommended path | Built-in alternative (custom + baseUrl) |
|---|---|---|
| Anthropic | model with anthropic_native.ts: native structured output and thinking control | https://api.anthropic.com/v1/, for trying it out only: Anthropic does not support it for production and it ignores response_format. No embeddings endpoint; pair it with another embedding |
| OpenAI | provider: "openai" (https://api.openai.com/v1, native schemas), or model with the OpenAI SDK | same |
| Azure OpenAI | provider: "azure_openai", baseUrl: "https://<resource>.openai.azure.com/openai/v1/", native schemas | same |
| Google Gemini | model with gemini_native.ts | https://generativelanguage.googleapis.com/v1beta/openai/ (beta per Google; JSON-schema output, images and embeddings work) |
| Vertex AI | model with gemini_native.ts (Application Default Credentials) | none listed here; see Google's docs on calling Vertex AI with OpenAI libraries |
| AWS Bedrock | model with bedrock.ts (Converse, every Bedrock model) | https://bedrock-runtime.<region>.amazonaws.com/openai/v1 with a Bedrock API key, for models whose card lists Chat Completions; or a LiteLLM proxy |
| Mistral | custom, or model with its SDK | https://api.mistral.ai/v1 |
| Groq | custom | https://api.groq.com/openai/v1 |
| DeepSeek | custom | https://api.deepseek.com |
| Together | custom | https://api.together.ai/v1 |
| Fireworks | custom | https://api.fireworks.ai/inference/v1 |
| OpenRouter | custom | https://openrouter.ai/api/v1 |
| Ollama | custom (chat and embeddings) | http://localhost:11434/v1 |
| vLLM | custom; schemaEnforced: true with guided decoding on | http://localhost:8000/v1 |
| LM Studio | custom (chat and embeddings) | http://localhost:1234/v1 |
| LiteLLM proxy | custom | http://localhost:4000 (the proxy's default port) |
| Any other gateway | custom | its OpenAI-compatible root |
Schema enforcement. Only OpenAI and Azure are claimed native by the
built-in client. Set schemaEnforced: true on an llm config only for an
endpoint you know enforces response_format json_schema; every other endpoint
is reported prompt, validated and repaired once. A host model makes the
same promise through mode on its reply.
Embeddings. Any vendor through embedder: a local model, a hosted API, a
typed array reply (Float32Array rows are accepted). Or through custom for
an OpenAI-format embeddings endpoint (Ollama, vLLM, LM Studio, Gemini's
compatible URL, a gateway). config.embedding.model and dim are the
database's identity: the lock compares the model, so a database built through
one provider keeps working under custom or a host embedder as long as the
model string and the dimension match.
Reranking. Any reranker through reranker, a hosted rerank API or a local
cross-encoder alike:
async function myReranker(query: string, docs: string[]): Promise<number[]> {
const res = await fetch(process.env.RERANK_URL!, {
method: "POST",
headers: { authorization: `Bearer ${process.env.RERANK_KEY}`, "content-type": "application/json" },
body: JSON.stringify({ query, documents: docs }),
});
const { results } = (await res.json()) as { results: { index: number }[] };
return results.map((r) => r.index); // the API's order, most relevant first
}Per-purpose routing. request.purpose lets one function send map_reduce
to a cheap model and vision_* to a vision model:
async function myModel(request: ModelRequest): Promise<ModelReply> {
if (request.purpose === "map_reduce") return cheapModel(request);
if (request.purpose.startsWith("vision_")) return visionModel(request);
return defaultModel(request);
}Batch mode: deferred model calls
A provider batch API answers within a day at about half the price. The engine
awaits every model call as it goes, so instead of blocking for hours your
model or embedder throws ModelDeferred: the request is accepted and its
answer comes later. During ingest the document parks as
status: "waiting_model" with everything stored so far kept, and
engine.resumeDocuments() asks the same requests again, under the same keys,
once you have the answers. Batching stays yours: the engine never keeps a
half-finished conversation, your store is the memory.
import { ModelDeferred, type ModelReply, type ModelRequest, requestKey } from "@promptev/context-engine";
async function myModel(request: ModelRequest): Promise<ModelReply> {
const key = requestKey(request);
const answer = await store.get(key); // your store, keyed by requestKey(request)
if (answer) return { text: answer.text, tokens: answer.tokens }; // the original counts
await batch.enqueue(key, request); // your provider's batch, same key
throw new ModelDeferred({ retryAfterMs: 3_600_000 }); // until a resume is useful
}
// A cron on any worker, as often as you like: parked documents whose retry time
// has passed are claimed one at a time, so two workers never ask you twice.
const report = await engine.resumeDocuments({ contentSource: (doc) => blobs.get(doc.documentId) });
for (const doc of report.documents) console.log(doc.documentId, doc.status); // "completed" | "failed" | "waiting_model"
for (const park of report.superseded) await batch.drop(park.requestKeys); // nobody asks for these againnew ModelDeferred({ retryAfterMs }) takes milliseconds (null or omitted
is no hint); a negative or non-finite value is refused at construction. A
stage never stops at its first deferral: it asks every request in the stage,
collects the deferrals, and parks once. The stages after chunking (chunk
embedding, structured extraction with its registry embedding, entity
extraction) run in the same pass, so a PDF takes at most two rounds, page
images then everything after chunking. With a 24-hour batch window that is up
to about 48 hours.
What parks and what does not.
| Call | On ModelDeferred |
|---|---|
| vision_pages, vision_image | every batch is still asked; the document parks at stage extract, storing no partial text, so no page is ever dropped |
| chunk embedding (kind: "document") | a deferred batch leaves its rows without a vector; parks at stage enrich |
| structured_extraction and its registry embedding | no repair re-ask on a deferral (the repair prompt would change the key); parks at stage enrich |
| entity_extraction | every batch is asked; if any deferred, nothing from the pass is stored, so a resume never writes a batch twice; parks at stage enrich |
| community_summaries and the community embedding (only with graph.communitySummaries on) | the document does not park: it completes, the source's communities are stored without summaries, and the next resumeDocuments fills them per source |
| the query embedding (kind: "query") | treated as an embedder outage: the vector leg is dropped, the text legs answer, onError fires. Answer it at once |
| a reranker | the fused order stands and onError fires (stage rerank), as on any reranker error |
| map_reduce, compute | refused: the host model deferred a {purpose} call. Deferral is only for ingest; a query needs its answer now. A caller is waiting, so there is nothing to park |
While a document is parked nothing is billed and nothing is reported: no
usage event, no onError, and the parked stage sends no done progress
event. Its stored chunks are still found by the text legs. getDocument and
every listing show status: "waiting_model" and a parked block beside
metaData: stage, parkedAt, retryAt, rounds, requestKeys,
lastSkipReason and lastSkipAt. The knowledge tool's document rows
carry the same two fields, spelled as on its wire (parked_at, retry_at,
request_keys, last_skip_reason, last_skip_at). resumeEmbeddings leaves a parked document alone
and reports skipped with skipReason: "waiting_model".
Resuming. resumeDocuments({ documentIds, contentSource }) resolves to a
ResumeReport. With no documentIds it sweeps every parked document whose
retryAt has passed (the database clock; retryAt is the park time plus
the smallest retryAfterMs any deferral gave, or the park time when none
did), oldest first, in pages of 100. Naming ids resumes those whether due or
not; one that is not parked is skipped without a word, so a resume is
idempotent. Each document is claimed by one compare-and-set, so a cron on
several workers never asks you twice for the same request. A document that
defers again parks again with rounds incremented; one that completes gets
the same finalize as an ingest, with one usage event.
contentSourceis for a document parked at stageextract: called with aParkedDocument(documentId,sourceId,externalId,name,contentSha256) it returns the original bytes as aUint8Array, sync or async, ornull. Without it, withnull, or when it throws (onErrorfires once with stagecontent_source), the document stays parked with reasonno_content_source. Bytes whose sha256 is not the recorded one leave it parked withcontent_mismatch; two files are never mixed. A parked stageenrichis rebuilt from the database and needs no bytes.ResumeReport.documentsis oneDocumentReportper document worked on (completed,failedorwaiting_modelagain);skippedis{ documentId, reason }per document left parked on this call, reasonno_content_source,content_mismatchorredaction_mismatch, listed on every call until a resume succeeds and shown bygetDocumentasparked.lastSkipReason/lastSkipAt;supersededis oneSupersededPark(documentId,requestKeys,supersededAt,reason) per set of request keys nobody will ask for again, reported once:reason: "superseded"for a park a re-ingest replaced,reason: "failed"for the keys a pass deferred before it failed outright (a real embedding failure after a deferral, say) or an expired park held; that document's own ingest or resume report carries them too, asabandonedRequestKeys. A document mid-ingest reports nothing until its row settles.communitiesResumedis the source ids whose community summaries this call filled, in whole or in part.- A re-ingest of a parked document wins. The old request keys are
reported once in
superseded, so you can drop those batch requests. An unchanged re-offer (the same bytes or text under the same arguments, an hourly folder sync say) is not a re-ingest: it leaves the park as it is, asks nothing, and returns thewaiting_modelreport. The permissions, name, description and metadata it declares are still applied to the row and its stored chunks, as on any unchanged re-ingest: the latest declaration wins, also over anupdateDocumentmade in between. A changed one that parks again asks the unchanged pages and chunks under their old keys, and those keys are left out ofsuperseded. A key that is reported can still be owed by another parked document that shares the request (the same page image, the same chunk text); dropping it costs that document one more round. An anonymous ingest (noexternalId, noname) that parks at the page reading stage is not deduplicated against a second anonymous upload of the same bytes; give documents an identity if that matters. - An edit made while a document is parked is kept.
updateDocumentworks on a parked document as on any other, and the row is what a resume reads: permissions,name,descriptionandmetaDataare the ones the row holds when the resume runs, never the ones the ingest call declared when the document parked. A principal removed while a document waits for a batch stays removed when it completes. A parked document shows the declaredmetaDatafrom the moment it parks, and a document deleted while parked stays deleted. ingest.maxParkSeconds(unset ornull, wait for ever) bounds the current round: a document whose latest park is older than that is finalizedfailedwithfailureReason: "model_deferred_expired"on the next resume, andonErrorfires then.
The request key. requestKey(request) is the sha256 hex of the canonical
JSON of {"v": 1, purpose, system, user, images, response_schema,
schema_name, json_mode, max_tokens, temperature, thinking_budget}, where
images is the sha256 of each image's bytes in order. It is the same in
TypeScript and Python and on every resume: a schema repair or a vision re-ask
is deterministic given the same earlier answers, so a follow-up request gets
the same key each time. The key covers exactly what your model receives.
The key covers the prompt text, so an engine upgrade that changes a prompt
changes the key of every request made with it: resume your parked documents
before you upgrade, or expect the upgraded engine to ask for them again
under new keys.
Embeddings are keyed PER TEXT, not per batch: embeddingRequestKey(text, {
kind, model }) with model your config.embedding.model, because a resume
re-embeds only the chunks still missing a vector and may pack them into
different batches. Store one vector per text key; answer a batch only when
every text in it is answered, and otherwise queue the missing ones and defer
the whole batch. Namespace your answers by model on your side.
Tokens. A resume re-asks every model call of the parked stage, so return a stored answer with its ORIGINAL token counts: model-call tokens are billed once, in the completing pass. Work a parked round already stored is never asked again and is billed once at completion: chunk vectors, registry vectors, and the page text of a round that parked after chunking.
Redaction, and what your store holds. For every text purpose the key is
computed from post-redaction text, the text your model saw. Vision is the
exception: a vision_pages or vision_image request, and its key, is built
from the original page images, which redaction cannot reach, so the batch
requests you store for vision hold unredacted document content. Delete stored
batch requests and answers once the document completes. Resume from the same
engine, or the same withRedaction view: a different policy or key would
send different text and change every key, so such a resume leaves the
document parked with redaction_mismatch. The same goes for the other port:
TypeScript and Python do not compute the same fingerprint for the same
policy, so a document parked under a redaction policy is resumed by the port
that parked it.
Community summaries (only with graph.communitySummaries on; with it off
this pass is skipped) are asked again per source, not per document. Every
resumeDocuments call, ids named or not, re-runs the summaries for each
source that has communities without one, under the same per-source lock a
graph ingest takes, lists the sources it completed in communitiesResumed,
and bills each as one usage event (detail.operation resume_communities).
A summaries call that failed outright at ingest is retried the same way; a
resume that fails or defers again is neither billed nor listed and is retried
next call. A reply that summarises only some of the communities it was asked
about keeps the ones that came back: the source is listed and billed for
those, stays pending, and the next call asks for the rest alone, in a smaller
request with its own key. Superseded community keys are not reported: answer by key, and
ignore a request that is never asked again. A round whose summaries were
answered but whose community embedding deferred asks the summaries key again.
The source's lock is held on a pooled client that sits idle while your model
answers and detection needs a second client meanwhile, so community resume
needs storage.poolMax of at least 2 (a pool of one is skipped with one
onError), and a database idle_in_transaction_session_timeout shorter than
that answer ends that source's resume with onError, and the source stays
pending.
Configure
ContextEngineConfig is the one settings object, built directly or from
CE_-prefixed environment variables with ContextEngineConfig.fromEnv()
(nested fields use __ and names are camel-cased, so CE_EMBEDDING__API_KEY
becomes embedding.apiKey; overrides are deep-merged on top). databaseUrl
and embedding are required; with a host model nothing else is. The
constructor validates it: half a Neo4j config, defaultMode: "graph" with
graph off, or a hash redaction rule with no key all throw at construction;
graph mode without extractionLlm throws when the engine is built without a
host model.
const config = new ContextEngineConfig({
databaseUrl: "postgresql://user:pass@localhost:5432/mydb",
embedding: { provider: "openai", model: "text-embedding-3-small", dim: 1536 },
graph: { enabled: true }, // Postgres-only graph mode: no neo4jUri
reranker: { candidates: 50 }, // the window a host reranker sees
search: { vectorFloor: "adaptive" }, // opt in, see Search
ingest: { maxRows: 500_000 },
storage: { poolMax: 20, poolIdleTimeoutMs: 30_000 },
secretKey: "...", // base64url, 32 bytes: tool-config encryption
});The fields that hold a secret (databaseUrl, secretKey,
redactionSecretKey, every apiKey, neo4jPassword) are a Secret on the
built config. Pass plain strings, in code or in the environment; read one back
with .get() or secretValue(field), both exported. JSON.stringify,
util.inspect, console.log, String() and a spread show "**********" in
their place, so a host that serialises a config to rebuild it later must dump
the secrets itself, and a field must never be templated into a header or a
URL. A Secret passed in is kept as is, so a config can be rebuilt from
another config's fields. A validation error never prints its input either.
embedding carries the identity the database records (provider, model,
dim; dim is detected on the first embed when unset) and the request
packing: maxBatchItems (256), maxBatchTokens (250,000) and
maxInputTokens (8,192) default to the OpenAI-compatible row and are
overridable for an endpoint whose limits differ; tokenizer ("auto" and
"exact" both count with cl100k, since js-tiktoken bundles its ranks;
"estimate" uses UTF-8 bytes / tokenBytesRatio 3.0 / safetyMargin 0.85);
maxConcurrency (4) embedding requests in flight per document.
llm and embedding both take maxRetries (absent or null resolves to 2;
0 disables; CE_LLM__MAX_RETRIES, CE_EMBEDDING__MAX_RETRIES,
CE_VISION_LLM__MAX_RETRIES, CE_GRAPH__EXTRACTION_LLM__MAX_RETRIES), applied
by the built-in client only. The call timeout is per attempt, so one call's
ceiling against a hung endpoint is (maxRetries + 1) x timeout + backoff,
about 12 minutes at the defaults (240 s x 3); structured extraction re-asks
once after an unusable reply (maxReasks, 1), and vision's batch attempts
multiply again. Set maxRetries: 0 on a latency-sensitive path. Every
optional field on these types is .nullable().optional() with no default, so
an absent field and an explicit null differ where the docs say so.
storage holds the deployment knobs: annExactThreshold (50,000; the vector
leg's exact/approximate crossover) and writeBatchRows (absent = 32; the
most chunk rows one write statement carries). A document's chunks and vectors
are written in statements of at most writeBatchRows rows, and each
embedding statement commits on its own. Writing a vector also writes its
index entry, so on a small database a statement carrying a few hundred
vectors can run past a statement_timeout. Lower the setting there, or for
wide embeddings; raise it on a large instance to save round trips; null
puts no cap on a statement. A statement that is cancelled loses only its own
rows: the document ends failed with every earlier vector stored, and
ingesting the same content again embeds only the chunks still missing one.
A caller who may see that many chunks or
fewer gets an exact sort of them, which returns the true top results and
costs time in proportion to the rows sorted: just under the threshold (48,378
visible chunks of 384 dimensions, on a laptop, measured with the Python
package, which runs the same query) the vector leg took 337 ms at the median
and 646 ms at the 95th percentile, against 18 ms for the index scan on the
same caller at recall 0.980. Lower the threshold to trade that last recall
for speed, raise it to keep the exact route for more callers, and measure at
your own dimension
(benchmarks/acl_recall/).
Also here are the pg.Pool settings poolMax (10),
poolIdleTimeoutMs (10,000) and poolConnectionTimeoutMs (30,000). pg has
no pre-ping, so a connection the database has dropped fails once and is
replaced. Reachable as env vars too, for example CE_STORAGE__POOL_MAX.
extraction: maxRenderPx (2000, the long edge of rendered pages and sent
images), visionMaxOutputTokens (16,000 per vision batch),
pagesPerVisionBatch (5), visionMissingPageRetry ("single_after_cut"),
ocrLanguage (detected when unset). ingest: the intake caps and resume
fence, under Ingest.
visionMissingPageRetry decides how the pages a vision batch did not answer
are asked for again. "single_after_cut": after a reply that reached
visionMaxOutputTokens, each missing page gets its own call, so a page the
model cannot stop writing about loses only itself. A reply that is merely
short is asked again in one call. "single": one page per call after any
incomplete reply. "batch": the missing pages always go back together, which
is the fewest calls. Raising visionMaxOutputTokens does not rescue a runaway
page, it only makes it longer.
Provision the database
npx context-engine migrate --database-url postgresql://user:pass@localhost:5432/mydb --dim 1536
# add --graph to also create entity/relationship/community tables--dim must match your embedding model's width (1536 for
text-embedding-3-small) and is locked on first migrate. Run it again on every
deploy: it is idempotent, and the migration history ships inside the package.
Also from code: runMigrate(databaseUrl, { dim: 1536, graph: false }).
migrate does not write-lock your tables: every index after the initial
schema is built with CREATE INDEX CONCURRENTLY, so ingest keeps writing
while the build runs; it prints one line to stderr per index it builds, and
SELECT * FROM pg_stat_progress_create_index shows how far along it is. A
failed migrate is safe to re-run, and re-running is the repair: every
statement is idempotent, and an interrupted concurrent build is dropped and
rebuilt by the next run. A lock_timeout or statement_timeout on the
connection applies to these builds; migrate will not override it.
Ingest
import { readFileSync } from "node:fs";
// Plain text
let report = await engine.ingest({
text: "Annual leave accrues at two days per month for every band-three employee.",
name: "leave-policy",
sourceId: "hr-handbook", // a namespace to scope search and stats
externalId: "leave-policy", // your own id; with sourceId it identifies the document on re-ingest
acl: ["group:hr"], // omit (or null) for an unrestricted document, visible to everyone
});
// A file: PDF/DOCX/PPTX/XLSX/CSV/HTML/EML/images/plain text, auto-detected
report = await engine.ingest({ content: readFileSync("policy.pdf"), filename: "policy.pdf", sourceId: "hr-handbook" });
console.log(report.totals); // { files: 1, failed: 0, units: 3, ... }
console.log(report.documents[0].status); // "completed" | "failed" | "skipped" | "embedding" | "waiting_model"Exactly one of file (a path, a Buffer, a multer-style { buffer, originalname }
or a handle with read()), content + filename, or text (skips
extraction). ingest() never throws for a bad document; check
report.documents[i].status. Construction opens nothing; the pool, embedder
and graph store are built on first use. Release them with aclose(), or let
await using do it at the end of a scope. A process-lifetime singleton can
skip it; anything building engines per job or per test should not.
- What a document is is decided in this order:
mimeType: "text/plain"(or any other mime type) when you pass it, then the filename's extension (.txtand.mdincluded; thenamestands in for it on atextingest), then the content. Only the last is a guess. Text is read as CSV only on strong evidence: three or more rows, the same number of commas (two or more) on at least four lines in five, and cells that are values, not sentences. Prose with commas in it stays plain text. PassmimeTypewhen you know better: a tabular type keeps a document out of graph mode. The content check is narrow on purpose. It does not recognise a two-line CSV or a header on its own, a two-column file, tab- or semicolon-separated text, or a CSV whose last column is free-text sentences. Give those a.csvname ormimeType: "text/csv".mimeTypeis accepted byengine.ingestand by the HTTP ingest route, as themime_typefield of the JSON body or of the multipart form. mode: "hybrid"(default) ormode: "graph"(needs the graph peers andgraph.enabled).extractStructured: trueruns a structured-extraction pass (document_type,structured_data) through the host model orconfig.llm;fieldHintsmakes it return exactly the hinted keys.- Scanned pages and images are transcribed to structured Markdown (tables
and headings preserved) by the host
modelorvisionLlm, through the same seam as every other call, so any model that reads images works; with neither, embedded text is used, withtesseract.jsas a fallback.sharpand@napi-rs/canvasonly add page rasterization. Born-digital PDFs keep their structure too withpdfjs-dist: the text layer is converted to Markdown locally, no model, no network. - A re-ingest only pays for what changed. Identical content is
deduplicated by hash and skipped; when the content did change, every chunk
whose text did not keeps the vector the previous version paid for, copied
inside the database, and only the changed chunks go to the embedder (matched
on a hash of the whitespace-normalised text, and only for the same embedding
provider/model/dimension). Two callers racing the same document are settled
by the database: the loser reports
skippedwithskipReason: "in_progress"and writes nothing, ACL updates included, so re-offer it once the winner is done. - A document with unreadable pages is read again. A document that
completed with
unreadablePages > 0is not treated as finished: ingesting the same bytes again extracts the whole file again, so a better model, a changed setting or a second try can recover the lost pages. It is chunked and embedded again only when the text came back different. A second read that lost more pages than the first is discarded and the stored text is kept. The read itself is paid for every time, so setingest.rereadUnreadablePages: falsewhen a scheduled sync keeps offering files with a page no model can read. - Chunks are budgeted to
min(2000, embedding.maxInputTokens), because a chunk is one embedding input. Changing provider or that cap re-chunks each document on its next ingest. - Two intake caps, both refusing rather than trimming.
ingest.maxFileBytes(256 MiB) is answered without reading the document (the file size for a path,Content-Lengthfor an upload).ingest.maxRows(unset) counts.csv,.tsvand.xlsxonly, since the extension is what routes a file to a tabular reader; the streaming reader's count is the authoritative one. Over either throwsIngestTooLarge, which the handlers map to 413 naming the knob.nulllifts either (an absent field takes the default); nothing is ever sampled or truncated.uploadLimitBytes(engine)hands the byte cap to your own upload middleware (multerlimits.fileSize, fastifybodyLimit). - Chunks are stored before they are embedded, and requests are packed by
tokens as well as by count against
embedding's limits. A single chunk overmaxInputTokensis never split: the document fails withfailureReason: "input_too_large"(the error behind it isInputTooLarge, which is whatembedChunksthrows when you call it yourself). With the built-in client, a size rejection halves the cap for that run, and a one-item batch still refused fails the document the same way (EmbeddingBatchRejected); a host embedder is never retried or halved. - A failure keeps every batch already paid for. The document sits at the
non-terminal status
embeddingwithgetDocument()'sembeddingProgressat{ done, total, tokens }, andawait engine.resumeEmbeddings(documentId)embeds only the rows still missing a vector. Called with no id it sweeps every resumable document, fenced byingest.resumeStaleSeconds(600,nulllifts) so a cron job cannot take over a run that is still moving; naming an id is never fenced. A partly embedded document is still findable on the text legs.EmbedBatchFailedcarries how far the run got. - A host model that answers later parks the document. A
modelorembedderthat throwsModelDeferredleaves the document at the non-terminal statuswaiting_model, andawait engine.resumeDocuments()asks again under the same request keys; see Batch mode. - Progress is visible.
onProgressfiresstarted/doneper stage and, on theembedstage, one"progress"event per completed batch carrying{ done, total, tokens }(ingest.progressBatches: falseturns those off). Switch on the state and ignore one you do not know. - A file that could not be read throws
ExtractionFailedwith the library's error as itscause;ExtraMissingError(a missing optional peer, with an install hint) andIngestTooLargepass through unwrapped.
Why a document failed
A failed document carries two answers. error is the raw String(exc), an
operator's copy: a provider body has carried an API key, a model name and a
quoted fragment of the document, so getDocument, listDocuments and
DocumentReport.error keep it and the model-facing doors never return it.
failureReason is one code from a fixed vocabulary, classified from the
error rather than its text (input_too_large, provider_rejected,
provider_unavailable, extraction_failed, chunk_set_changed,
model_deferred_expired, embedding_dimension_mismatch, unknown),
and failureMessage is the one fixed sentence for that code.
embedding_dimension_mismatch is an embedder (yours or the built-in
client's) returning vectors of a different width than embedding.dim or than
this database's vector columns: check the model and the dimension the
database was migrated with. The knowledge
tool's get_doc, get_docs and list, the MCP tool behind them and
GET /documents/{id} on every router return the code and the sentence; a
router mounted with principals: () => TRUSTED still gets error. A later
success clears both.
Spreadsheets and CSVs
CSV/XLSX documents are indexed in hybrid mode, even in a graph-mode corpus
(modeUsed: "hybrid" with a fixed modeReason): rows are emitted as RFC 4180
CSV, streamed rather than materialised, and each sheet gets a leading chunk
with meta.chunk_kind === "schema" naming the columns, their inferred types
and a few sample values, pointing the model at compute for arithmetic. Leave
ingest.maxDocumentTextChars unset unless you never compute: compute
rebuilds its tables from the stored text, so a cut body is a spreadsheet that
quietly answers with fewer rows than it has. Setting it is recorded
(meta_data.text_truncated), and compute and discover refuse a truncated
document by name.
Search and access control
const result = await engine.search("how much annual leave do I get", {
sourceIds: ["hr-handbook"],
documentIds: null, // optionally pin to specific documents WITHIN the sources
principals: ["group:hr"], // TRUSTED = trusted/internal (skips ACL); [] = anonymous
topK: 10,
compressToTokens: 2000, // optional: trim weak hits to fit a prompt budget
});
for (const hit of result.hits) {
console.log(hit.documentName, hit.idx, hit.score, hit.chunkText.slice(0, 80));
}Full-text, trigram and vector legs, fused with reciprocal rank fusion
(fusion: k 20, per-leg weights); mode: "graph" adds the graph leg.
documentIds intersects sourceIds inside the same SQL predicate every leg
shares, so a foreign id can never widen a search. Every hit carries idx, its
position inside the document (the knowledge tool reports it as chunk_idx,
which get_chunks takes as start). score is the fusion score even when a
reranker decided the order; the list order is the ranking.
Long questions and the trigram leg. The trigram leg compares the query's
trigrams with every chunk's, so its cost grows with the length of the query.
Measured with the Python package, which runs the same queries, on 100,646
chunks with questions of 16 words at the median, a hybrid search at the
defaults took about 2.6 s at the median (2,563 ms for a caller who sees 10% of
the corpus), nearly all of it in the trigram leg. search.trgmMaxQueryWords
skips that leg for a query of more words than the setting. At 8 the median
search fell to about 0.11 s (113 ms), because 95% of those questions then skip
the leg. At 16 it barely moved (2,160 ms): 53% of the questions are 16 words
or fewer and still run it. The results change with the setting: at 8, 92% to
95% of the default's results are still in the top ten. The setting is unset by
default, so the trigram leg always runs and nothing changes until you set it.
It is your choice: set it when your queries are whole sentences and search
time matters, and leave it unset when they are names, codes and identifiers,
which are short and are what the leg is for. The times are one laptop, one
query at a time, with the query embedding not counted, and other work was
running on that machine during these runs: they show the size of the effect
and are not a latency benchmark. Conditions and every
row:
benchmarks/acl_recall/.
principals is a trichotomy, and the dangerous value is the one that looks
like an absence. TRUSTED (import it) disables filtering on purpose. [] is
an anonymous caller and matches only documents with no ACL. A list matches
unrestricted documents plus any whose ACL overlaps it. null or an omitted
principals is accepted for compatibility and means TRUSTED too, with a
DeprecationWarning (CE_PRINCIPALS_NULL); a later release refuses it,
because it is what every accessor degrades to (user.groups || null, a
missing key, user ? user.groups : null), which makes the unauthenticated
path the maximum-privilege path. Never pass null because a request had no
authenticated user:
import { TRUSTED } from "@promptev/context-engine";
await engine.search(q, { principals: TRUSTED }); // skip ACL, deliberately
await engine.search(q, { principals: [] }); // anonymous
await engine.search(q, { principals: user.groups }); // scopedOn remote surfaces principals are injection-only: derive them server-side
from the session, never from a request body. Filing a document under an ACL
the caller does not hold is a 403 on ingest and on PATCH alike. Do not add a
WHERE of your own on top, and do not cache a SearchResult across users.
Access control that survives the vector index. Filtering an approximate
index with a WHERE clause is a post-filter: the index picks candidates
first, your ACL discards them second, and the query returns short. So this
library puts the permission check inside every leg's query, and for the vector
leg it counts the rows a caller may see and sorts them exactly when there are
50,000 or fewer (storage.annExactThreshold). Measured with the Python
package, which runs the same queries, on a 100,000-document corpus with
synthetic team-shaped permissions: a caller who may see 10% or 1% of the
documents gets from every leg exactly the top ten that a corpus holding only
their documents returns, while the same search filtered afterwards returns
nothing for 64% of hybrid searches at 10% and 99% at 1%, and recovers 0.42 and
0.02 of the results when it fetches ten times more first. Above the threshold
the vector leg is an approximate index scan: it scored 0.979 to 0.988 with
most of the corpus visible, and 0.652 for a caller who sees 1% when that path
was forced at this size. It needs pgvector 0.8 or newer; below that,
ACL-filtered vector search silently loses recall. Method, conditions and data:
benchmarks/acl_recall/. Measure your own corpus,
offline, read-only, no embedder called:
npx context-engine check-acl-exposure --database-url postgresql://...Thresholds follow the query's shape
pg_trgm's similarity threshold moves with the query (query-shape.ts, the
same rules in both ports): 0.45 for a two-character query down to 0.15 for a
long sentence, measured in characters for spaceless scripts (Chinese,
Japanese, Thai, Khmer, Burmese, Lao, Tibetan) and in words for everything
else. search: { trgmLimit: 0.3 } pins it to one value.
The vector leg has a matching floor, off by default: search: { vectorFloor: "adaptive" }
drops candidates under a minimum cosine similarity chosen from the query (0.45
for one or two words, 0.40 for a code or an acronym, 0.25 for a sentence)
before fusion. Rows under it leave that leg only, can still arrive through
full-text or trigram, and the leg can return fewer than limit with nothing
backfilled. Measure before turning it on: those numbers came from one
embedding model, and cosine bands differ enough between models that an
absolute floor can empty the vector leg for short queries on another. An
unknown value is refused by the schema rather than read as off.
Keyword search settings
The full-text leg matches with Postgres text search over the simple
configuration, which keeps every word, stopwords included. Seven settings under
search shape it; the defaults are the first value of each.
| Setting | Values | What it trades |
|---|---|---|
| lexicalMatch | "all", "any" | "all" parses the query with websearch_to_tsquery, so a chunk must contain every word: precise for keyword queries, but a question such as "what is the notice period for contractors" usually matches nothing, because no chunk contains "what" and "is" and every other word. "any" matches a chunk containing any of the query's words after dropping stopwords, and leaves the ordering to the rank function. It finds more candidates, so it suits question-shaped queries and leans harder on ranking and fusion. |
| lexicalRank | "ts_rank_cd", "ts_rank", "bm25" | ts_rank_cd (cover density) rewards query words that occur close together. ts_rank counts how often they occur, wherever they are in the chunk. bm25 (experimental, opt-in with context-engine bm25-enable) is Okapi BM25 computed in SQL: rare words count for more, repeating a word saturates, and long chunks are normalised. See BM25 ranking. |
| bm25 | { k1: 1.2, b: 0.75, idfScope: "source" } | BM25's parameters, read only when lexicalRank is "bm25". Every field is optional. |
| lexicalNormalization | 0 to 63 | The rank function's normalization bitmask, passed to Postgres unchanged. 0 ignores chunk length. 1 divides by 1 + log(length), 2 by length, 4 by the mean distance between matches (ts_rank_cd only), 8 by unique words, 16 by 1 + log(unique words), 32 maps the rank into 0 to 1. Length normalization stops a long chunk outranking a short one just by repeating a word. |
| lexicalStopwords | null or an array | The words "any" mode drops from the query. null uses the built-in English list, LEXICAL_STOPWORDS; [] keeps every word. They are parsed by the same text-search configuration as the query, so case does not matter. A query made only of stopwords, with no phrase or -term, keeps its words. Ignored in "all" mode. |
| trgmMaxQueryWords | null or a whole number of 1 or more | Skips the trigram leg for a query of more than this many words, exactly as a fusion weight of 0 would for that query: it is not sent or fused, and it is absent from usage.legHits and timings. A word is a run of characters that are not Unicode White_Space, so punctuation stays part of its word and a query in a script written without spaces (Chinese, Japanese, Thai) counts as one word. The trigram leg is the slowest on long queries and helps most on short ones (names, codes, misspellings). null runs it for every query. |
| lexicalCandidates | null or a whole number of 1 or more | How many candidates the full-text and trigram legs each hand to fusion. null uses the depth every leg uses: ten times the requested result count, between 50 and 500. A larger value is capped at 500. The vector and graph legs are unaffected. |
const config = new ContextEngineConfig({
// ...
search: { lexicalMatch: "any", lexicalRank: "ts_rank_cd", lexicalNormalization: 1 },
});BM25 ranking (experimental)
lexicalRank: "bm25" orders the keyword leg's matches with Okapi BM25, in
plain SQL inside your Postgres. It matches exactly the chunks the other ranks
match (lexicalMatch still decides that, "all" or "any"), applies the
same permission check to every row before it is scored, and only changes the
order.
const config = new ContextEngineConfig({
// ...
search: { lexicalRank: "bm25", bm25: { k1: 1.2, b: 0.75, idfScope: "source" } },
});| bm25 field | Default | What it does |
|---|---|---|
| k1 | 1.2 | Term-frequency saturation, 0 or more. 0 counts only whether a word is present; larger values let repetition matter longer. |
| b | 0.75 | Length normalisation, 0 to 1. 1 fully penalises a chunk for being longer than average, 0 ignores length. |
| idfScope | "source" | Where the corpus statistics (how rare a word is, the average chunk length) come from. "source": the chunk's own source, so one source's documents never change another source's scores. "global": every source together; other sources' (other tenants') documents shift the scores. "document": the chunks of the chunk's own document, counted at query time. |
Each query word in a chunk adds
ln(1 + (N - df + 0.5) / (df + 0.5)) * tf * (k1 + 1) / (tf + k1 * (1 - b + b * dl / avgdl)),
where N is the number of chunks in the statistics scope, df how many of them
contain the word, tf how often it occurs in the chunk, dl the chunk's length
and avgdl the average length. Lengths and counts come from the stored
tsvector (Postgres keeps at most 256 positions per word per chunk). Negated
terms (-term) never add to a score. lexicalNormalization is ignored; b
does that job. The scores are computed by SQL functions the migration installs,
so the Python and TypeScript packages return the same numbers on the same
database.
Opt-in, per database. Nothing changes for a database that does not use BM25: the migration only adds two small tables and some SQL functions, and never touches the chunk table. To use it, run once:
npx context-engine bm25-enable --database-url postgresql://...or
await engine.enableBm25(). It installs four triggers on the chunk table (taking its lock briefly, with a 2-second timeout and retries) and counts the chunks already stored, one source at a time, without blocking ingest. Running it again recounts. Until it is enabled and has finished counting, alexicalRank: "bm25"search ranks withts_rank_cd, reports aLexicalRankFallbacktoonError, and says why inusage.lexical({ rank: "bm25", rankUsed: "ts_rank_cd", reason: ... }):bm25_not_enabledbeforebm25-enablehas run,bm25_stats_pendingwhile it is still counting (or if one of its triggers has since been dropped; running it again repairs that), andbm25_schema_missingbeforemigrate. It never ranks with incomplete numbers.idfScope: "document"counts live and needs no enabling.bm25-disable(engine.disableBm25()) removes the triggers and the statistics.Statistics, and nobody waits. How many chunks contain a word is counted at query time from the full-text index, for the query's words and the sources being scored only (
idfScope: "global"counts across the whole corpus and pays for it). The chunk count and total length per source are kept by the triggers in an insert-only table, one small row per chunk write, in the same transaction. No row is shared between writers and no lock is taken, so two ingests into the same source never wait on each other, and an unchanged re-ingest nets to zero.Compaction runs on its own. After a write commits, a source holding more than
search.bm25.compactAfterDeltasrows (default 1000) is folded into one row. It never makes the write wait: when another maintenance call holds the source, it is skipped until a later write.nullturns it off; then runnpx context-engine bm25-compact(engine.compactBm25Stats()) now and then. Recounting (bm25-enable) and compacting take a per-source lock that only maintenance uses, so any number of them may overlap each other and ingest (several replicas enabling at startup, a scheduled compaction) and the numbers stay exact. Chunks deleted outside this library's own methods (for example a cascade from your own SQL) are counted at once, but their source is only compacted at its next write through the library. While BM25 is not enabled the write path never looks for compaction work: it checks the enabled flag at most once every 5 seconds.Statistics include documents the caller may not see. How rare a word is is a corpus number, as in every search engine. No content or id is exposed and a hidden chunk is never scored or returned, but its words move other chunks' scores.
idfScope: "document"limits that to the chunk's own document.Measured: not a default, and not an improvement to default search. On the five public datasets above,
lexicalRank: "bm25"withlexicalMatch: "any"raised hybrid nDCG@10 on MuSiQue (+0.039), 2WikiMultihopQA (+0.016) and MultiHop-RAG (+0.062), made no measurable difference on ArguAna (+0.004) and lowered it on CQADupStack (-0.043). With the default"all"matching it changes almost nothing on question-shaped queries. The tables are inbenchmarks/retrieval/results/bm25.md. Measure on your own queries before turning it on.Cost. Measured on 1,000-chunk writes: with BM25 not enabled, ingest is unchanged; enabled, an insert took about 10 ms more (184 against 172 ms) and a delete about 7 ms more. A
bm25query scores every matching chunk and counts each query word in the index, so its cost grows with how common the words are; it was measured only on corpora of 7,000 to 12,000 chunks.Another BM25 engine. To rank with an external BM25 service or another Postgres extension instead, plug it in as a custom retrieval leg (see Plug in your own retrieval); its results are re-filtered through the same permission check before fusion.
Quoted phrases and -term keep their meaning in both modes. Take the query
"notice period" contractor -draft:
- in
"all"mode it matches a chunk that contains the exact phrase "notice period" AND the word "contractor", and does not contain the word "draft"; - in
"any"mode it matches a chunk that contains the phrase OR the word "contractor" (or both), and does not contain the word "draft".
What the defaults rest on. The keyword and trigram defaults were measured
with the retrieval benchmark in this repository's benchmarks/retrieval/, on
five public datasets (MuSiQue, 2WikiMultihopQA, MultiHop-RAG, ArguAna and the
CQADupStack English forum), 300 queries each, with the BAAI/bge-small-en-v1.5
embedding model; the tables are in
benchmarks/retrieval/results/lexical-defaults.md and
benchmarks/retrieval/results/lexical-defaults-2.md.
- The fusion constant
fusion.kis20. Against60,20raised hybrid nDCG@10 by 0.039 to 0.049 on four of five datasets, with no measurable change on the fifth;100lowered it on four. - The trigram leg's weight in
fusion.weightsis0.4.0.4
