llm-scorer
v0.1.3
Published
A tiny JSONL LLM judge
Readme
llm-scorer
Judge supplied model answers against reference answers with an OpenAI-compatible chat endpoint. It uses OpenRouter Auto Router by default.
Quick start
Provide two aligned JSONL files. The answer key contains question and its
accepted expectedAnswers array:
{"question":"What color is the sky?","expectedAnswers":["blue","Blue"]}The proposed-answer file contains the same question and the model's answer:
{"question":"What color is the sky?","answer":"The sky appears blue."}Run the scorer with an OpenRouter API key, an answer key, and proposed answers:
LLM_SCORER_QAFILE=sample-answer-key.jsonl \
LLM_SCORER_ANSWERS_FILE=sample-proposed-answers.jsonl \
OPENROUTER_API_KEY=your_openrouter_api_key \
npx llm-scorerRun llm-scorer with no arguments, or with --help, to see the available
command-line options. The same configuration can be supplied as flags:
llm-scorer --file sample-answer-key.jsonl --answers sample-proposed-answers.jsonl --model openrouter/auto --output results.jsonlThe default endpoint is https://openrouter.ai/api/v1/chat/completions, and
the default model is openrouter/auto. Results are appended to results.jsonl
in the current directory unless you configure a different path.
Configuration
| Variable | Default | Description |
| --- | --- | --- |
| LLM_SCORER_QAFILE | Required | Path to the answer-key JSONL file. |
| LLM_SCORER_ANSWERS_FILE | Required | Path to the proposed-answers JSONL file. |
| LLM_SCORER_URL | https://openrouter.ai/api/v1/chat/completions | OpenAI-compatible chat-completions endpoint. |
| LLM_SCORER_MODEL | openrouter/auto | Model used to judge the supplied answers. |
| LLM_SCORER_RESULTS_FILE | results.jsonl | File to which per-question results are appended. Parent directories are created automatically. |
| OPENROUTER_API_KEY | — | Bearer token used automatically for OpenRouter endpoints. |
| LLM_SCORER_API_KEY | — | Bearer token for non-OpenRouter endpoints. |
For example, to select both a model and a results location:
LLM_SCORER_QAFILE=sample-answer-key.jsonl \
LLM_SCORER_ANSWERS_FILE=sample-proposed-answers.jsonl \
LLM_SCORER_MODEL=anthropic/claude-sonnet-4 \
LLM_SCORER_RESULTS_FILE=output/my-results.jsonl \
OPENROUTER_API_KEY=your_openrouter_api_key \
npx llm-scorerScoring
Every supplied answer is evaluated semantically by LLM_SCORER_MODEL against
the rubric in JUDGE.md; the result includes the rationale and resolved model.
No request is made to generate an answer to the question.
Each output line is a JSON result with scores and modelIdentifier. When the
endpoint resolves a router alias, such as OpenRouter Auto, that resolved model
identifier is recorded.
Edit JUDGE.md to customize the grading rubric.
Programmatic use
import { runJudge } from "llm-scorer";
const result = await runJudge({
file: "sample-answer-key.jsonl",
answersFile: "sample-proposed-answers.jsonl",
output: "output/my-results.jsonl"
});