jev-reranker
v0.1.5
Published
Rerank JSON search results with TypeSafe AI's Jev
Readme
Jev Reranker
Choose the search results worth passing to your LLM.
jev-reranker uses TypeSafe AI's Jev to rerank retrieved
documents, remove candidates that contain no usable evidence, or extract query-specific passages
for your LLM.
It reads a JSON array from stdin and writes a JSON array to stdout. Choose the field that contains the text; IDs, source paths, retrieval scores, and other metadata pass through unchanged. Use it after BM25, vector search, or any command that emits a JSON array of candidate objects.
search or vector database -> JSON candidates -> Jev Reranker -> context for your LLMInstall
Requires Node.js 14 or later on macOS, Linux, or Windows (x64 or Arm64).
Install the CLI from npm:
npm install --global jev-rerankerYou can also run it without a global installation:
npx jev-reranker --helpCreate an API key in the TypeSafe dashboard, then export it:
export TYPESAFE_API_KEY="your-api-key"Try It
Pipe an array of candidate documents into the CLI:
printf '%s\n' '[{"text":"Build artifacts are cached locally."},{"text":"Access tokens expire after one hour."}]' \
| jev-reranker --query "How long do access tokens last?"The CLI returns the same objects in best-first order, with a rerankScore from 0 to 1 added to
each one. Higher values mean Jev considers the result more relevant to the query.
Bring Your Own Results
Suppose a search command returns objects shaped like this:
[
{
"id": "auth-guide",
"title": "Authentication",
"body": "Access tokens expire after one hour.",
"distance": 0.18,
"source": "/docs/auth.md"
}
]Tell jev-reranker which fields to use:
search-command --json \
| jev-reranker \
--query "How long do access tokens last?" \
--text-field body \
--context-field title \
--top 5body is the document text. title is prepended as context, which helps short chunks that are
ambiguous on their own. distance, id, and source pass through unchanged. Existing retrieval
scores do not affect Jev's judgment or the output order.
Pass more candidates than you plan to keep and let --top trim the output. Jev can move a relevant
document that search ranked low to the top. On FiQA, reranking BM25's top 100 instead of its top 30
raised nDCG@10 from 0.36 to 0.40. Every candidate is scored before --top applies, so each
additional 30 candidates adds one API request.
Choose a Mode
| Mode | Use it to | Output |
| --- | --- | --- |
| rerank (default) | Put relevant candidates first. | Original objects sorted by rerankScore. |
| filter | Remove candidates that provide no usable evidence. | Retained objects in input order, with evidenceScore. |
| compress | Send shorter passages to your LLM. | Retained objects in input order, with compressedText. |
These modes run separately. To rank and then filter or compress, pipe one invocation into another with the same query. Each invocation makes its own API requests.
Rerank and filter make one request per 30 candidates. A batch too large for one request is split, which adds requests. Compression scores every sentence or line, so long candidates need more requests, and each request repeats the full text and selected context of the documents it judges. See TypeSafe's current Jev pricing.
Keep the useful evidence
A result can mention the right topic without providing an answer. Filter mode asks Jev whether each candidate contains concrete evidence, including partial answers, conditions, and exceptions.
search-command --json \
| jev-reranker --query "When can I request a refund?" --mode filterCandidates with evidenceScore >= 0.5 survive. Their input order stays intact, so you can keep
an upstream ranking you already trust. --top 5 returns the first five survivors; it does not
sort them by evidence score. If none qualify, the result is [].
Extract shorter passages
search-command --json \
| jev-reranker --query "When can I request a refund?" --mode compressCompress mode splits the selected text into sentences and lines, then asks Jev which units to keep. Jev sees the full source text and selected context when judging each unit, with instructions to retain relevant conditions, exceptions, and references needed to understand the evidence.
compressedText contains retained passages in source order, with surrounding whitespace removed.
Adjacent retained text stays together; nonadjacent passages are separated by newlines.
The original text and metadata remain available. The example below shows the output shape:
{
"text": "Refunds are available within 30 days. Opened items are excluded. Our offices close at six.",
"source": "/docs/refunds.md",
"compressedText": "Refunds are available within 30 days. Opened items are excluded."
}Pass compressedText to your downstream LLM to reduce its context. Documents with no selected units
are omitted.
Sentence/line extraction is intended for prose. For code, tables, and unusual formatting,
whole-document filter mode may work better. Selected context fields help Jev interpret the text
but are not copied into compressedText.
JSON Contract
Stdin must contain one JSON array. Each item must be an object, and the field selected by
--text-field must be a string. The field defaults to text.
The CLI preserves the original fields and writes the selected mode's output field: rerankScore,
evidenceScore, or compressedText.
Existing values under the selected mode's output field are replaced. Other modes' output fields pass through unchanged. Text and context fields cannot use the current mode's output field name.
Equal rerank scores retain their input order. Filter and compress preserve input order throughout.
--top limits output objects after the selected mode has processed all candidates.
You may repeat --context-field. Present string values are prepended in flag order, while missing
and null values are skipped. Other context value types are rejected.
An empty array returns [] without reading the API key or making a request. Compress also returns
[] without a request when every selected text is empty or contains only whitespace.
What Leaves Your Machine
The query, model name, selected context values, and selected text are sent directly to TypeSafe's
System One API. Other object fields, including the source score, stay local. The API key is read
only from TYPESAFE_API_KEY; there is no command-line key option.
Results are buffered until every batch succeeds, so a failed request leaves stdout empty instead of producing a partial JSON document. Error messages can end up in logs, so the CLI never writes document text, request headers, or your API key into them. When the API rejects a request, the API's own short error message is shown as returned.
Options
| Option | Default | Description |
| --- | --- | --- |
| --query <string> | Required | Query used to judge relevance. |
| --text-field <name> | text | Object field containing the text to score. |
| --context-field <name> | None | Context field to prepend. May be repeated. |
| --mode <rerank\|filter\|compress> | rerank | How to select or order context. |
| --top <n> | All results | Maximum output objects, at least 1. |
| Option | Default | Description |
| --- | --- | --- |
| --threshold <number> | 0.5 | Minimum score, from 0 to 1, to keep a document or unit. Filter and compress only. |
| --model <name> | jev-latest | Jev model route. |
| --timeout-ms <n> | 10000 | Timeout for each HTTP attempt, in milliseconds. |
Agent Skill
An Agent Skill teaches coding assistants to pipe search results through
jev-reranker before answering. Install it for your assistant:
npx jev-reranker skills install --claude-code # this project
npx jev-reranker skills install --claude-code --global # every project
npx jev-reranker skills install --codex--path <dir> installs it anywhere else. The skill matches the installed CLI version, so run the
command again after upgrading.
Benchmarks
On three BEIR datasets, reranking BM25's top 30 raised nDCG@10 by 0.06 to 0.13:
| Dataset | BM25 order | jev-reranker |
| --- | --- | --- |
| SciFact | 0.68 | 0.76–0.77 |
| NFCorpus | 0.27 | 0.33 |
| FiQA | 0.24 | 0.36–0.37 |
Each query cost less than $0.001. BENCHMARKS.md has the setup, the effect of candidate depth, run-to-run variation, and the limits of these numbers.
Background
What Retrieval Still Hasn't Decided covers why these three modes are separate judgments rather than stages of one pipeline, what each one was measured against, and the deduplication mode that is not here.
