@davothebigafro/eval
v0.0.0
Published
Evaluation runner, judge-client interfaces, test-case schemas, prompt targets, caching, and custom metric primitives.
Readme
@davothebigafro/eval
Evaluation runner, judge-client interfaces, test-case schemas, prompt targets, caching, and custom metric primitives.
Install
pnpm add @davothebigafro/eval --filter <your-package>Inside this workspace, keep local dependencies on workspace:*. For unreleased local installs, use the tarball workflow in the root contributing guide.
Minimal Usage
import { evaluate, type LLMMetric } from '@davothebigafro/eval';
const exactMatch: LLMMetric = {
name: 'exact_match',
threshold: 1,
strictMode: false,
includeReason: true,
requiredParams: ['actualOutput', 'expectedOutput'],
async measure(testCase) {
const passed = testCase.actualOutput === testCase.expectedOutput;
return {
score: passed ? 1 : 0,
reason: passed ? 'Matched.' : 'Did not match.',
breakdown: [],
thresholdMet: passed,
};
},
};
const result = await evaluate({
judge,
testCases,
metrics: [exactMatch],
});The default judge cache is fresh and in-memory per run. Pass a JudgeCache through evaluate({ cache }) for durable or shared caching.
