@securityvoid/strands-evals-typescript
v0.1.0
Published
Evaluation framework for Strands Agents and LLM applications
Maintainers
Readme
@securityvoid/strands-evals-typescript
Community TypeScript port of the Python strands-agents-evals package. Evaluation framework for Strands Agents and LLM applications.
Not an official Strands package. Published under
@securityvoidso it can ship without the@strands-agentsnpm org; the CLI binary remainsstrands-evals.
Install
npm install @securityvoid/strands-evals-typescript @strands-agents/sdk zodRequires Node.js >=20. Optional extras (Langfuse, OpenSearch, LangChain, OTEL exporter) are declared as optional peer dependencies — install only what you use.
Quick Start
import { Case, Experiment } from '@securityvoid/strands-evals-typescript'
import { CorrectnessEvaluator } from '@securityvoid/strands-evals-typescript/evaluators'
const experiment = new Experiment({
cases: [new Case({ input: 'What is 2+2?', expectedOutput: '4' })],
evaluators: [new CorrectnessEvaluator()],
})
const report = await experiment.runEvaluations(async (c) => {
// your agent / task
return { output: '4' }
})
console.log(report)CLI
npx strands-evals --helpThe strands-evals binary supports run, validate, report, diagnose, generate, and fetch subcommands (parity with the Python CLI).
Default judge model
global.anthropic.claude-sonnet-4-6 (aligned with the Python package). LLM-as-a-judge evaluators use this unless you pass a different model.
Contributing / agents
- Contributing: CONTRIBUTING.md
- Specs: docs/specs/
- AI assistants: start at AGENTS.md; architecture map in docs/architecture.md
- Port ledger: docs/specs/typescript-port/PARITY.md
