@agentskit/eval
v0.6.2
Published
Agent evaluation and benchmarking for AgentsKit.
Maintainers
Readme
@agentskit/eval
Profile: concise-package
Measure agent quality with numbers, not vibes — ship with confidence.
Tags: ai · agents · llm · agentskit · ai-agents · eval · evaluation · benchmarking · testing · ci-cd · llm-testing
Verified proof
- Package metadata and tests live under
packages/eval/. - Package guide: https://www.agentskit.io/docs/packages/eval
- Stability map: docs/STABILITY.md
How this fits the ecosystem
@agentskit/eval is the quality layer: replay agent runs, compare outputs, snapshot behavior, and catch regressions before users do.
- AgentsKit: compose it with the other packages in this repo to build agents from small, swappable parts.
- Registry: look for ready agents and templates that already use this layer at registry.agentskit.io.
- Playbook: learn the production patterns behind this layer at playbook.agentskit.io.
- AKOS: run the same concepts with enterprise deployment, governance, and observability at akos.agentskit.io.
Docs: package guide · agent handoff
Why eval
- Replace "it seemed to work" with real metrics — accuracy, per-case latency, token cost, and pass/fail for every test case in a single result object
- CI/CD ready — exit codes reflect suite results; gate deployments on accuracy thresholds so regressions never reach production
- Flexible assertions — substring matching or full control with a custom
(result) => booleanpredicate per case - Provider-agnostic — the
agentclosure can wrap any async boundary:createRuntime, a custom controller, or an HTTP endpoint - Failure isolation — malformed responses, agent failures, and assertion failures become failed cases without aborting the suite
Install
npm install @agentskit/evalQuick example
import { runEval } from '@agentskit/eval'
import { createRuntime } from '@agentskit/runtime'
import { anthropic } from '@agentskit/adapters'
const runtime = createRuntime({
adapter: anthropic({ apiKey: process.env.ANTHROPIC_API_KEY, model: 'claude-sonnet-4-6' }),
})
const result = await runEval({
agent: async (input) => {
const r = await runtime.run(input)
return r.content
},
suite: {
name: 'qa-baseline',
cases: [
{ input: 'What is 2+2?', expected: '4' },
{ input: 'Capital of France?', expected: 'Paris' },
{ input: 'Is TypeScript a superset of JavaScript?', expected: (r) => r.toLowerCase().includes('yes') },
],
},
})
console.log(`Accuracy: ${(result.accuracy * 100).toFixed(1)}%`)
console.log(`Passed: ${result.passed}/${result.totalCases}`)Features
runEval({ agent, suite })— run a named test suite against any agent function- Result:
{ accuracy, passed, totalCases, cases[] }— per-case latency and outcome - Assertion modes: exact match,
includes, custom predicate - CI exit codes — non-zero on failure for pipeline gating
- Pair with
@agentskit/observabilityto trace failed cases
Deterministic replay
Import recording, replay, cassette serialization, time travel, and comparison APIs from @agentskit/eval/replay. That entry is safe for Node, browsers, Expo, and React Native; browser and native package conditions exclude filesystem code.
Node applications that persist cassettes to disk should import saveCassette and loadCassette from the explicit Node-only subpath:
import { createCassette } from '@agentskit/eval/replay'
import { saveCassette, loadCassette } from '@agentskit/eval/replay/io'
await saveCassette('./fixtures/session.json', createCassette())
const cassette = await loadCassette('./fixtures/session.json')The original Node exports from @agentskit/eval/replay remain available for compatibility. Browser and native builds retain those export names so types and runtime stay aligned, but calls reject with a Node-only diagnostic. Those hosts should combine serializeCassette or parseCassette with their own storage APIs.
@agentskit/eval/snapshot and @agentskit/eval/ci are Node-oriented subpaths because they write files and inspect CI environment variables. Pure comparison helpers remain callable anywhere their entry can be bundled, but browser applications should not depend on their filesystem workflows.
Replay boundaries snapshot cassettes, requests, chunks, dates, and plain metadata. Mutating a live request or a replayed chunk does not rewrite the recorded cassette.
Braintrust lifecycle
@agentskit/eval/braintrust computes every score locally. When BRAINTRUST_API_KEY or options.apiKey is present, it lazily initializes Braintrust, awaits each experiment log, flushes the experiment, and then summarizes it. Non-fatal SDK failures are returned as bounded result.warnings; local cases and summaries remain available.
Custom scorer results are validated at runtime. A malformed name or score outside [0, 1] becomes an isolated scorer_error rather than corrupting the aggregate.
Ecosystem
| Package | Role | |---------|------| | @agentskit/runtime | Typical agent under test | | @agentskit/core | Stable I/O contracts for eval harnesses | | @agentskit/observability | Traces for failed cases |
Contributors
License
MIT — see LICENSE.
Docs
Maturity and compatibility
- Stability: beta — see docs/STABILITY.md
- Node.js 20+ and TypeScript strict mode
- Published as
@agentskit/eval
Contributing
See CONTRIBUTING.md and the monorepo LICENSE.
