arc-skill-eval
v0.26.1
Published
Pi-native library and CLI that runs Anthropic-standard skill evals (evals/evals.json) with LLM-judged + script assertions.
Maintainers
Readme
arc-skill-eval
Pi-native library and CLI for running skill evals. Authoring format follows Anthropic's published evals/evals.json standard. The eval methodology — layered grading, small starter suites that grow from real failures, the with-skill / without-skill comparison as the load-bearing signal — is directly inspired by OpenAI's Testing Agent Skills Systematically with Evals (Kundel & Chua, Jan 2026). The runtime philosophy ("an LLM, a loop, and enough tokens") borrows from Ampcode's How to Build an Agent and Mihail Eric's The Emperor Has No Clothes. See Inspiration & credits for the full attribution.
What it does
Given a skill that ships SKILL.md and a sibling evals/evals.json, arc-skill-eval:
- discovers every
SKILL.md+evals/evals.jsonpair under a repo. - materializes each case's optional
files/into a temp workspace. - runs the case through the Pi SDK with the skill attached.
- grades the outputs — string assertions via an LLM-judge,
file-exists/regex-match/json-validvia deterministic scripts. - writes per-case
assistant.md+outputs/+timing.json+grading.json+ observability artifacts under<skill>/evals-runs/<runId>/. - tracks model, thinking level, token usage, estimated cost, context-window size, and context percentage used.
- records tool-call counts, skill reads, external calls, MCP-looking tool calls, and the context/tool manifest exposed to the model.
- optionally compares each case against a no-skill baseline with
--compare.
Assertion grading mirrors OpenAI's layered approach (deterministic checks first, model-assisted rubric for prose) and emits artifacts in Anthropic's published grading.json shape.
Input format
<skill-dir>/evals/evals.json:
{
"skill_name": "arc-conventional-commits",
"evals": [
{
"id": 1,
"prompt": "Set up semantic-release in this repo.",
"expected_output": "semantic-release configured with the Conventional Commits preset.",
"files": ["files/clean-repo"],
"assertions": [
{ "type": "file-exists", "path": ".releaserc.json" },
{ "type": "regex-match", "pattern": "conventionalcommits", "target": { "file": ".releaserc.json" } },
"The response summarizes the semantic-release plugins it installed."
]
}
]
}Each case may also set "sandbox": "just-bash" to run inside an isolated virtual bash environment instead of the default temp-workspace runner ("none"). A --sandbox CLI flag overrides this per run. In just-bash mode the agent's bash tool executes in an in-process virtual shell with a filesystem rooted at the case workspace, so command execution needs no host shell and the repository working tree is never touched.
just-bash ships core unix builtins; npm, npx, and git get deterministic no-op success mocks by default. Override them per case with sandboxMocks to return specific output, exit codes, and file effects:
{
"id": "install-deps",
"prompt": "Install dependencies.",
"sandbox": "just-bash",
"sandboxMocks": [
{
"command": "npm",
"stdout": "added 1 package\n",
"exitCode": 0,
"files": [{ "path": "node_modules/.installed", "content": "ok" }]
}
]
}Requirements
- Node.js ≥ 20
- Pi installed and configured with at least one provider API key (Anthropic, OpenAI, Google/Gemini, Mistral, xAI, etc.). The skill's assistant runs via
@mariozechner/pi-coding-agent.
Install
From a local checkout
npm install
npm run build
npm link
arc-skill-eval --helpFrom a published package
npm install --global arc-skill-eval
arc-skill-eval --help
arc-skill-eval run "$(arc-skill-eval bundled hello-world)"Usage
# Scaffold a starter eval suite next to a SKILL.md
arc-skill-eval create ./skills/my-skill
# Preview the generated evals.json without writing it
arc-skill-eval create ./skills/my-skill --dry-run
# Review a human-readable summary of generated cases/assertions
arc-skill-eval create ./skills/my-skill --dry-run --summary
# Ask a configured model to propose richer starter cases without writing files
arc-skill-eval create ./skills/my-skill --guided --dry-run --summary
# Interactively accept, skip, or edit proposed cases/assertions before writing
arc-skill-eval create ./skills/my-skill --guided --interactive
# Run every eval in every discovered skill under the current repo
arc-skill-eval run .
# Run one skill
arc-skill-eval run ./skills/arc-conventional-commits
# Run one case inside one skill
arc-skill-eval run ./skills/arc-conventional-commits --case 1
# Pin the skill runner model and LLM-judge model
arc-skill-eval run ./skills/arc-conventional-commits \
--model openai-codex/gpt-5.5:medium \
--judge-model mistral/ministral-8b-latest
# Create a tiny eval-owned Pi config/runtime directory
arc-skill-eval init-runtime ./.arc-skill-eval/pi-agent \
--provider ollama-cloud \
--model gpt-oss:20b
# Use an eval-owned Pi config/runtime directory
arc-skill-eval run ./skills/hello-world \
--agent-dir ./.arc-skill-eval/pi-agent \
--model ollama-cloud/gpt-oss:20b \
--judge-model ollama-cloud/gpt-oss:20b
# Generate a static HTML review report and feedback template from run artifacts
arc-skill-eval review ./skills/hello-world/evals-runs/<runId>
# Propose eval improvements from review feedback without writing files
arc-skill-eval improve --from-feedback ./skills/hello-world/evals-runs/<runId>/feedback.json \
--dry-run --summary
# Retarget output to a different workspace root
arc-skill-eval run . --output-dir ./evals-runs
# Machine-readable JSON
arc-skill-eval run . --json
# Opt into with_skill vs without_skill comparison
arc-skill-eval run . --compare
# Group artifacts under an iteration bucket
arc-skill-eval run . --iteration 1
# Add explicit distractor/conflict skills to the model context
arc-skill-eval run ./skills/arc-conventional-commits \
--compare \
--extra-skill ./skills/release-please \
--iteration conflict-1
# Opt into normal Pi ambient resources such as configured extensions/tools
# while recording the resulting loadout in context-manifest.json
arc-skill-eval run ./skills/arc-conventional-commits \
--context-mode ambient \
--iteration ambient-1Recommended first dogfood run
The companion andysolomon/arc-skills repo ships a dogfood suite for arc-creating-evals, the meta-skill that authors eval suites for other skills. After cloning both repos locally, this is the best end-to-end smoke test:
arc-skill-eval run /path/to/arc-skills/arc-creating-evals \
--case execution-golden-path-file-skill \
--model openai-codex/gpt-5.5:medium \
--judge-model openai-codex/gpt-5.5:mediumFor the with-skill / without-skill signal:
arc-skill-eval run /path/to/arc-skills/arc-creating-evals \
--case execution-golden-path-file-skill \
--compare \
--iteration dogfood-1 \
--model openai-codex/gpt-5.5:medium \
--judge-model openai-codex/gpt-5.5:mediumA recent dogfood run passed the golden-path case and showed a positive +16.7% with-skill delta after tightening the suite to assert behavior unique to arc-creating-evals.
Create starter evals
Generate a valid starter suite for a skill directory:
arc-skill-eval create ./skills/my-skillThe command reads SKILL.md frontmatter, writes evals/evals.json, and includes three starter cases:
trigger-explicitexecution-golden-pathadjacent-negative
When obvious output artifacts are mentioned in SKILL.md, such as plan.md or report.json, the execution case also gets deterministic file-exists and json-valid assertions. When likely input files are mentioned, such as notes/input.md, requirements.md, prd.md, issue.md, or task.md, the execution case gets seeded fixture inputs under evals/files/starter-inputs/. The adjacent-negative case is domain-aware for common skill types like eval authoring, planning, releases, docs, and auth/webhooks, with a generic fallback. Use --dry-run to print the proposed JSON without writing files, --summary to print a human-readable review of generated cases/assertions, and --force to overwrite an existing evals/evals.json.
Use deterministic create first when the skill has concrete file, JSON, or command-line effects. It is fast, repeatable, CI-friendly, and never spends model tokens. Use create --guided when the hardest part is deciding what to test: conceptual interview skills, planning/review skills, routing skills with subtle adjacent negatives, or skills where success is mostly semantic. Guided mode asks the configured Pi model to design a richer proposal using the bundled skills/arc-creating-evals/SKILL.md procedure, then validates the returned evals.json with the same loader used by run before printing or writing anything. You can pin the designer with --model <provider/model[:thinking]>, use --agent-dir <path> for eval-owned model/auth lookup, or pass --authoring-skill <path> to test a different eval-authoring skill.
Use interactive guided mode to review the proposed suite before it is written:
arc-skill-eval create ./skills/my-skill --guided --interactiveThe lightweight prompt flow presents the rationale, cases, fixture inputs, and assertions; lets you include/skip cases and assertions; and lets you edit case prompts, expected output, and judge/regex assertion text. Existing overwrite protections still apply unless --force is supplied.
For example, a conceptual grill-me skill that conducts a relentless interview may not create files at all. Its suite should lean on judge assertions such as "asks direct follow-up questions about assumptions and tradeoffs" plus adjacent negatives that should not trigger the skill, rather than fake file-exists checks. That makes the eval measure the behavior the skill actually promises.
Prefer behavior-focused assertions
Write assertions against observable behavior and artifacts, not incidental wording. Brittle wording checks fail when a correct assistant paraphrases, changes a heading, or omits a phrase the skill never promised.
Prefer:
{ "type": "file-exists", "path": ".releaserc.json" },
{ "type": "regex-match", "pattern": "conventionalcommits", "target": { "file": ".releaserc.json" } },
"The response names semantic-release and explains that it configured release automation for this repository."Avoid unless the words are truly the product requirement:
"The response says exactly: Phase 1 — detection complete."Exact wording is appropriate for user-facing contracts such as a required commit message, CLI output, email subject, or safety disclaimer. When wording is required, make it explicit in expected_output and use a deterministic regex-match or exact output assertion so failures explain the missing text directly.
The positional <skill-dir-or-repo> for run is resolved as:
- a skill directory if it contains
evals/evals.json, - otherwise a repo whose tree is walked for SKILL.md + evals/evals.json pairs.
Audit skill quality
Run deterministic skill-authoring checks without invoking a model:
arc-skill-eval audit ./skills
arc-skill-eval audit ./skills/my-skill --json
arc-skill-eval audit ./skills --output skill-audit.mdaudit reports frontmatter issues, long descriptions, SKILL.md sprawl, missing evals/evals.json, broken local markdown reference links, trigger-heavy descriptions on user-invoked skills, and likely duplicate skill families. It exits successfully by default so it can be used as a report generator; use the finding counts in JSON output if CI needs custom failure thresholds.
Review reports
Turn a run directory into a static review bundle:
arc-skill-eval review ./skills/hello-world/evals-runs/<runId>This writes review.html and feedback.json into the run directory. Use --output <dir> to write elsewhere and --force to overwrite an existing report. Compare runs are rendered with with_skill and without_skill variants side-by-side.
A practical create-run-review-improve loop looks like this:
arc-skill-eval create ./skills/my-skill --guided --interactive
arc-skill-eval run ./skills/my-skill --case execution-golden-path
arc-skill-eval run ./skills/my-skill --compare --iteration dogfood-1
arc-skill-eval review ./skills/my-skill/evals-runs/iteration-dogfood-1/<runId>
arc-skill-eval improve \
--from-feedback ./skills/my-skill/evals-runs/iteration-dogfood-1/<runId>/feedback.json \
--dry-run --summaryUse review.html to inspect assistant output, grading evidence, artifacts, and with/without-skill deltas. Capture human notes in feedback.json; feedback-driven improvement can then turn those notes into a focused plan for changing the skill, tightening assertions, or adding cases.
Improve from feedback
The command reads human notes plus failing assertion summaries and proposes prompt, assertion, fixture, or adjacent-negative changes with rationale. It does not change files unless you pass --apply. Applied changes annotate matching eval cases with validated improvement metadata so the suite remains loadable by run.
Browse runs interactively
Open an interactive terminal run browser (an Ink TUI) over the artifacts under evals-runs/:
arc-skill-eval browse ./skills/arc-conventional-commits # one skill
arc-skill-eval browse . # whole repoIt renders a lazygit-style four-panel layout — Skills, Cases, Assertions, Runs — with the selected case's prompt, grading evidence, metrics, and with/without-skill comparison in the main pane. It reads the same per-case grading.json / timing.json artifacts that run emits, so no extra setup is needed.
Navigation:
See the Keybindings reference for the full keymap — it's generated from src/tui/keymap.ts, the same source the in-TUI ? overlay renders from, so the two can't drift. Highlights: Tab/1–4 panels, j/k move, →/l/↵ enter the detail pane, [/] cycle case mode, v raw grading.json, / filter, s sort, c pin baseline, r/R run, n new case, ? help, q quit.
Runs and authoring happen in-process, without leaving the TUI:
r/Rrun evals for the selection in a live run console (spinner, per-case progress, pass/fail summary); on completion the affected skill reloads in place and your selection is restored.Radds--compare(with_skillvswithout_skill).Escaborts an in-flight run;↵reloads and closes when it's done.oruns with custom flags (--model,--iteration,--extra-skill…). This is the one path that still uses a child process — the binary isarc-skill-eval(must be onPATH); override withARC_SKILL_EVAL_BIN.nscaffolds a new eval case intoevals/evals.json;frecords afeedback.jsonnote for the selected case (consumed byimprove).
Display options:
--no-baselinehides thewithout_skillcomparison rows in the detail pane (handy when you only ran the skill variant).- The TUI is capability-aware: truecolor hex degrades to 16-color ANSI on low-color terminals, and block/box glyphs (bars, status ticks, accent bar) fall back to ASCII off a UTF-8 locale. Force the fallbacks with
NO_COLOR/FORCE_COLOR=0(no color) orARC_TUI_ASCII=1(ASCII glyphs).
Model options:
--model <provider/model[:thinking]>pins the skill runner model instead of using Pi's configured default. Example:openai-codex/gpt-5.5:medium.--judge-model <provider/model[:thinking]>pins the model used for LLM-judged string assertions. Deterministic assertions do not use the judge.--agent-dir <path>points Pi settings, model registry, and auth lookup at an eval-owned agent directory instead of the normal~/.pi/agentdirectory.- When no model flags are supplied,
arc-skill-evalinherits Pi's default provider/model/thinking level from the effective Pi agent settings.
Export results to Laminar Evaluations (optional)
run --laminar additionally reports the run to Laminar's Evaluations view: one evaluation per run variant (with_skill / without_skill), one scored datapoint per case, grouped by skill name so variants can be compared side by side. It is off by default and entirely optional — local evals-runs/ artifacts remain the canonical record, and an export failure never fails the run.
LMNR_PROJECT_API_KEY=lmnr_... arc-skill-eval run . --laminarThe run summary prints a direct dashboard link per evaluation. Each datapoint carries numeric scores (pass_rate, passed, failed, total_tokens, cost_usd, duration_ms, tool_calls) and an output with the grading summary, per-assertion verdicts (assertion text, pass/fail, short evidence quote), and local artifact paths.
LMNR_PROJECT_API_KEYis required when--laminaris set; the command fails fast (before any case runs) and names the key if it is missing.LMNR_BASE_URLis optional;LMNR_PROJECT_NAMEoptionally overrides the evaluation group name (default: the skill name).- The Laminar Node SDK (
@lmnr-ai/lmnr) is an optional dependency, dynamically imported only when the flag is enabled — installs that never use Laminar don't pull it in. If it's missing when enabled, the run reports a clear error naming the package. - Exports carry grading verdicts, metrics, and artifact paths only (never full assistant text, prompts, or file contents). The
benchmark.jsondelta remains local.
See docs/concepts/artifacts → External observability for the full local-artifact → Laminar evaluation mapping.
Eval-owned Pi runtime
Use --agent-dir when you want reproducible team or CI runs without depending on personal Pi defaults:
arc-skill-eval run ./skills/hello-world \
--agent-dir ./.arc-skill-eval/pi-agent \
--model ollama-cloud/gpt-oss:20b \
--judge-model ollama-cloud/gpt-oss:20bCreate one with:
arc-skill-eval init-runtime ./.arc-skill-eval/pi-agent \
--provider ollama-cloud \
--model gpt-oss:20bUse --force to intentionally overwrite existing runtime files.
A minimal eval-owned runtime contains just:
.arc-skill-eval/pi-agent/
├── models.json
└── settings.jsonThe runner and default LLM judge both use this directory for Pi models.json, settings.json, and auth.json lookup when --agent-dir is supplied. run preflights this directory before executing cases and reports missing models.json, settings.json, provider/model entries, or required API-key environment variables once with an init-runtime remediation. Secrets should still be referenced by environment variable name, for example "apiKey": "OLLAMA_API_KEY", rather than committed as literal values.
Ollama / low-cost cloud and local runs
arc-skill-eval inherits model support from Pi. Ollama Cloud is a useful low-cost provider lane for smoke tests. A verified working example is:
arc-skill-eval run ./skills/hello-world \
--model ollama-cloud/gpt-oss:20b \
--judge-model ollama-cloud/gpt-oss:20bA recent run with ollama-cloud/gpt-oss:20b passed 2/3 hello-world cases. The failed case was model behavior on an ambiguous prompt, not provider failure: the model asked which name to use instead of defaulting to Hello, world!.
Pi can also be configured through Ollama's integration for local or proxied cloud models:
# Let Ollama install/configure Pi and launch an interactive session
ollama launch pi
# Configure Pi for Ollama without launching
ollama launch pi --config
# Example cloud model launch through Ollama
ollama launch pi --model qwen3.5:cloudAfter Pi lists Ollama models, use the same provider/model pinning flags:
arc-skill-eval run ./skills/hello-world \
--model ollama/qwen3.5:cloud \
--judge-model ollama/qwen3.5:cloudFor direct Ollama Cloud access, set OLLAMA_API_KEY and add an ollama-cloud provider to Pi's models.json:
{
"providers": {
"ollama-cloud": {
"baseUrl": "https://ollama.com/v1",
"api": "openai-completions",
"apiKey": "OLLAMA_API_KEY",
"models": [
{ "id": "gpt-oss:20b" },
{ "id": "ministral-3:3b" },
{ "id": "gemma3:4b" }
]
}
}
}For local Ollama setup, add an Ollama-compatible provider to ~/.pi/agent/models.json using http://localhost:11434/v1 and set defaultProvider / defaultModel in ~/.pi/agent/settings.json. Local models do not require OLLAMA_API_KEY.
For the runtime roadmap — a tiny eval-owned Pi config vs a future custom agent — see docs/agent-runtime-strategy.md.
Context options:
--extra-skill <path>can be repeated to add explicit skill directories orSKILL.mdfiles as distractor/conflict context. In--compare,with_skillreceives the target + extras, whilewithout_skillreceives extras only.--context-mode isolatedis the default: no ambient Pi skills, extensions, prompt templates, themes, or context files are loaded.--context-mode ambientopts into normal Pi ambient resources so extension tools/MCP-like tools and other configured resources can enter the context. The resolved loadout is recorded incontext-manifest.json.--sandbox none|just-bashselects the execution isolation for every selected case, overriding each case's ownsandboxfield.none(default) uses the temp-workspace runner;just-bashroutes the agent'sbashtool through an in-process virtual shell (filesystem rooted at the case workspace) so commands run without the host shell and never touch the repo tree. Generated files are still captured underoutputs/.npm/npx/gitresolve to deterministic mocks (no-op success by default, configurable per case viasandboxMocks).
Exit code: 0 when every case has no failing assertions, 1 otherwise.
Output layout
For each default single-variant run:
<skillDir>/evals-runs/<runId>/
├── eval-<case-id>/
│ ├── assistant.md # final assistant response text
│ ├── outputs/ # files produced by the run
│ ├── timing.json # duration, model, thinking, token/cost/context metrics
│ ├── grading.json # per-assertion passed + evidence
│ ├── trace.json # normalized runtime trace + raw telemetry refs
│ ├── tool-summary.json # tool calls, errors, skill reads, external/MCP activity
│ └── context-manifest.json # skills/tools/context exposed to the modelUse --iteration <name> to group artifacts under <skillDir>/evals-runs/iteration-<name>/<runId>/; for example --iteration 1 writes to iteration-1/<runId>/.
With --compare, each case writes isolated variant artifacts and the skill run root includes benchmark.json:
<skillDir>/evals-runs/<runId>/
├── benchmark.json # with_skill vs without_skill aggregate
├── eval-<case-id>/
│ ├── with_skill/
│ │ ├── assistant.md
│ │ ├── outputs/
│ │ ├── timing.json
│ │ ├── grading.json
│ │ ├── trace.json
│ │ ├── tool-summary.json
│ │ └── context-manifest.json
│ └── without_skill/
│ ├── assistant.md
│ ├── outputs/
│ ├── timing.json
│ ├── grading.json
│ ├── trace.json
│ ├── tool-summary.json
│ └── context-manifest.jsontiming.json includes runner observability:
{
"total_tokens": 12345,
"duration_ms": 50123,
"model": { "provider": "anthropic", "id": "claude-opus-4-5", "thinking": "medium" },
"thinking_level": "medium",
"token_usage": {
"input_tokens": 10000,
"output_tokens": 2000,
"cache_read_tokens": 300,
"cache_write_tokens": 45,
"total_tokens": 12345
},
"estimated_cost_usd": 0.1234,
"context_window_tokens": 200000,
"context_window_used_percent": 6.2
}tool-summary.json highlights behavior-level observability:
{
"tool_call_count": 8,
"tool_error_count": 0,
"tool_calls_by_name": { "read": 2, "bash": 3, "write": 2, "edit": 1 },
"skill_read_count": 1,
"skill_reads_by_name": { "arc-conventional-commits": 1 },
"external_call_count": 0,
"mcp_tool_call_count": 0
}context-manifest.json records the run loadout so skill/tool conflicts can be diagnosed:
{
"runtime": "pi",
"mode": "isolated",
"attached_skills": [{ "name": "arc-conventional-commits", "path": ".../SKILL.md", "role": "target" }],
"available_tools": [{ "name": "bash", "source": "builtin" }],
"active_tools": ["read", "bash", "edit", "write"],
"mcp_tools": [],
"mcp_servers": [],
"ambient": { "extensions": false, "skills": false, "prompt_templates": false, "themes": false, "context_files": false }
}grading.json per the Anthropic format:
{
"case_id": "1",
"assertion_results": [
{ "text": "file-exists: .releaserc.json", "passed": true, "evidence": "Found .releaserc.json (182 bytes)", "assertion": { "type": "file-exists", "path": ".releaserc.json" } },
{ "text": "The response summarizes the semantic-release plugins it installed.", "passed": true, "evidence": "\"installs @semantic-release/commit-analyzer + release-notes-generator\"", "assertion": "The response summarizes the semantic-release plugins it installed." }
],
"judge_model": { "provider": "mistral", "id": "ministral-8b-latest" },
"summary": { "passed": 2, "failed": 0, "total": 2, "pass_rate": 1.0 }
}judge_model records the LLM-judge that graded the prose assertions; it's omitted for cases with only deterministic checks. browse surfaces it per case. Precedence: the --judge-model selection, else the model that ran the case (which is known to be authenticated), else { "provider": "mistral", "id": "ministral-8b-latest" } as a last resort. Note that when the judge defaults to the runner's model, the model grades its own output — pin --judge-model to a different model when that bias matters.
Authoring an eval suite for a skill
Use the bundled arc-creating-evals skill in skills/arc-creating-evals/. It interviews you across Anthropic's four success dimensions (outcome, process, style, efficiency) and emits evals/evals.json + fixtures. Install the skill into your agent's skills directory (.claude/skills/ or the equivalent for your tool) — see skills/README.md for the recipe.
Docs
docs/skill-eval-authoring-debrief.md— detailed research debrief and playbook for creating evals for skills, including thearc-skillsmastery roadmap.docs/agent-runtime-strategy.md— runtime strategy for Pi-backed evals, tiny eval-owned Pi config, Ollama Cloud, and a possible future custom agent.docs/skill-creator-parity-plan.md— user stories and implementation plan for Claude skill-creator parity features.docs/skill-creator-parity-progress.txt— checkbox tracker for the skill-creator parity roadmap.docs/create-dogfood-report.md— findings from dry-runningcreateagainst realarc-skills.docs/evals-json-pivot.md— direction, milestone log, and what stays vs what was deprecated.docs/domain-model.md— runtime + grading entities.
Shipped since the MVP
The slim MVP of the pivot to the Anthropic format has since grown both of its planned follow-ups:
- Cross-iteration comparison — pin a baseline run with
cinbrowseto diff iterations against it, andRruns--compare(with_skillvswithout_skill) in place. - Human-review
feedback.json— written byreview(and byfinbrowse), consumed byimprove --from-feedback.
Remaining ideas live in docs/evals-json-pivot.md and ROADMAP.md.
Inspiration & credits
arc-skill-eval exists because three pieces of writing made it clear what to build, in what shape, and with what philosophy. Each one shaped a different layer:
- The eval methodology — every workflow choice in the grader and the suite-growth advice in the docs — comes from OpenAI's Testing Agent Skills Systematically with Evals by Dominik Kundel and Gabriel Chua (January 22, 2026). The framing of an eval as "a prompt → a captured run (trace + artifacts) → a small set of checks → a score you can compare over time", the layered-grading recipe (fast deterministic checks first, then model-assisted rubric), the multi-category success metrics (outcome / process / style / efficiency), and the guidance that "a small set of 10–20 prompts is enough to surface regressions" — these are OpenAI's, transposed onto Anthropic's published format.
- The eval format — the on-disk
evals/evals.jsonshape, the per-casegrading.json, the aggregatebenchmark.json, and thewith_skill/without_skillcomparison — comes from Anthropic's documented skill-eval methodology. The framework consumes Anthropic's format so a skill author can take theirevals.jsonto any compatible runner. - The runtime philosophy — the bias toward a small, legible runtime that's not afraid to call itself a loop — owes a debt to two posts that demystified the agentic harness:
- Thorsten Ball's How to Build an Agent (Ampcode, April 15, 2025) — "It's an LLM, a loop, and enough tokens" — and the demonstration that a useful code-editing agent fits in a few hundred lines.
- Mihail Eric's The Emperor Has No Clothes: How to Code Claude Code in 200 Lines of Code (January 2026), which makes the same point at the level of agent harnesses: the core is a tool registry, an inner loop, and a parser. Production complexity is engineering, not architecture.
If you read only one of those before authoring an eval, read OpenAI's. If you read only one before extending the framework, read either of the harness pieces.
