@velrim/scoring
v0.1.0
Published
Velrim scoring/matching math (per-field P/R/F1, ECE, AUROC, Brier, risk-coverage) over the 3-state golden format. Pure TS, ESM, zero runtime deps. The same code velrim-eval and the published reliability curves run.
Readme
@velrim/scoring
The scoring and matching math behind velrim-eval
and Velrim's published reliability curves. Pure, deterministic, zero-runtime-dependency
TypeScript over the 3-state golden format (present / null / missing):
- per-field precision / recall / F1 (3-state confusion cells that score fabrication and omission distinctly)
- Expected Calibration Error (ECE) over 15 equal-mass bins
- Brier score
- AUROC (rank-sum / Mann–Whitney U, average-rank tie handling)
- the risk–coverage (selective-prediction) curve
scoreAgainstGolden(pred, gold)— the single entry point that integrates all of the above
This package exists so the math lives once. The Velrim API, the velrim-eval CLI, and the
published reliability curves all run this same code — when you re-score a golden set locally,
there is no second implementation to drift from.
Usage
import { scoreAgainstGolden, type GoldenDoc, type ScoringField } from '@velrim/scoring';
const gold: GoldenDoc = {
docClass: 'invoice',
fields: { '/total': { state: 'present', value: 1240.5 }, '/discount': { state: 'null' } },
};
const pred: Record<string, ScoringField> = {
'/total': { state: 'present', value: 1240.5, confidence: 0.91 },
'/discount': { state: 'null' },
};
const { perField, ece, auroc, riskCoverage } = scoreAgainstGolden(pred, gold);ScoringField is structural — { state; value?; confidence? } — so any extraction output
that carries a state and a confidence scores directly, no wrapper and no cast.
Matching semantics (0.1.0 amendments)
Two amendments landed in 0.1.0, pre-registered and applied identically to every system
scored with this code (see CHANGELOG.md for the full rationale):
- Absent-equivalence (FD-8). For the correctness label, a predicted
nullormissingmatches a goldennullormissing: an explicit-nullabstention is not penalized against an omitted-key golden, or vice versa. A goldenpresentstill requires a predictedpresentplus a value match. Per-field P/R/F1 cells were already keyed onpresentvs not-presentand are unchanged. - Value normalization (FD-10).
normalizeValue(value, kind)withkind: 'currency' | 'date' | 'text'normalizes both sides of a value match: currency to a canonical minimal decimal string ("$1,880.00"→"1880"), dates to ISO-8601YYYY-MM-DD(slash dates read in US MM/DD order, a pre-registered rule; two-digit years <50 → 20xx, else 19xx), text to trimmed/whitespace-collapsed/lowercased form. Unparseable currency/date values fall back to text folding. Strict matching stays the default: only leaves listed inoptions.normalizersnormalize.
// Strict column (unchanged default) vs normalized column (per-leaf opt-in):
const strict = scoreAgainstGolden(pred, gold);
const normalized = scoreAgainstGolden(pred, gold, {
normalizers: { '/total': 'currency', '/date': 'date', '/vendor': 'text' },
});The per-field pointer→kind tables are not part of this package; evaluations freeze them in their published analysis plans and pass them in.
ESM-only, "type": "module", Node ≥ 20.
