npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

arc-skill-eval

v0.26.1

Published

Pi-native library and CLI that runs Anthropic-standard skill evals (evals/evals.json) with LLM-judged + script assertions.

Readme

arc-skill-eval

Pi-native library and CLI for running skill evals. Authoring format follows Anthropic's published evals/evals.json standard. The eval methodology — layered grading, small starter suites that grow from real failures, the with-skill / without-skill comparison as the load-bearing signal — is directly inspired by OpenAI's Testing Agent Skills Systematically with Evals (Kundel & Chua, Jan 2026). The runtime philosophy ("an LLM, a loop, and enough tokens") borrows from Ampcode's How to Build an Agent and Mihail Eric's The Emperor Has No Clothes. See Inspiration & credits for the full attribution.

What it does

Given a skill that ships SKILL.md and a sibling evals/evals.json, arc-skill-eval:

  1. discovers every SKILL.md + evals/evals.json pair under a repo.
  2. materializes each case's optional files/ into a temp workspace.
  3. runs the case through the Pi SDK with the skill attached.
  4. grades the outputs — string assertions via an LLM-judge, file-exists / regex-match / json-valid via deterministic scripts.
  5. writes per-case assistant.md + outputs/ + timing.json + grading.json + observability artifacts under <skill>/evals-runs/<runId>/.
  6. tracks model, thinking level, token usage, estimated cost, context-window size, and context percentage used.
  7. records tool-call counts, skill reads, external calls, MCP-looking tool calls, and the context/tool manifest exposed to the model.
  8. optionally compares each case against a no-skill baseline with --compare.

Assertion grading mirrors OpenAI's layered approach (deterministic checks first, model-assisted rubric for prose) and emits artifacts in Anthropic's published grading.json shape.

Input format

<skill-dir>/evals/evals.json:

{
  "skill_name": "arc-conventional-commits",
  "evals": [
    {
      "id": 1,
      "prompt": "Set up semantic-release in this repo.",
      "expected_output": "semantic-release configured with the Conventional Commits preset.",
      "files": ["files/clean-repo"],
      "assertions": [
        { "type": "file-exists", "path": ".releaserc.json" },
        { "type": "regex-match", "pattern": "conventionalcommits", "target": { "file": ".releaserc.json" } },
        "The response summarizes the semantic-release plugins it installed."
      ]
    }
  ]
}

Each case may also set "sandbox": "just-bash" to run inside an isolated virtual bash environment instead of the default temp-workspace runner ("none"). A --sandbox CLI flag overrides this per run. In just-bash mode the agent's bash tool executes in an in-process virtual shell with a filesystem rooted at the case workspace, so command execution needs no host shell and the repository working tree is never touched.

just-bash ships core unix builtins; npm, npx, and git get deterministic no-op success mocks by default. Override them per case with sandboxMocks to return specific output, exit codes, and file effects:

{
  "id": "install-deps",
  "prompt": "Install dependencies.",
  "sandbox": "just-bash",
  "sandboxMocks": [
    {
      "command": "npm",
      "stdout": "added 1 package\n",
      "exitCode": 0,
      "files": [{ "path": "node_modules/.installed", "content": "ok" }]
    }
  ]
}

Requirements

  • Node.js ≥ 20
  • Pi installed and configured with at least one provider API key (Anthropic, OpenAI, Google/Gemini, Mistral, xAI, etc.). The skill's assistant runs via @mariozechner/pi-coding-agent.

Install

From a local checkout

npm install
npm run build
npm link
arc-skill-eval --help

From a published package

npm install --global arc-skill-eval
arc-skill-eval --help
arc-skill-eval run "$(arc-skill-eval bundled hello-world)"

Usage

# Scaffold a starter eval suite next to a SKILL.md
arc-skill-eval create ./skills/my-skill

# Preview the generated evals.json without writing it
arc-skill-eval create ./skills/my-skill --dry-run

# Review a human-readable summary of generated cases/assertions
arc-skill-eval create ./skills/my-skill --dry-run --summary

# Ask a configured model to propose richer starter cases without writing files
arc-skill-eval create ./skills/my-skill --guided --dry-run --summary

# Interactively accept, skip, or edit proposed cases/assertions before writing
arc-skill-eval create ./skills/my-skill --guided --interactive

# Run every eval in every discovered skill under the current repo
arc-skill-eval run .

# Run one skill
arc-skill-eval run ./skills/arc-conventional-commits

# Run one case inside one skill
arc-skill-eval run ./skills/arc-conventional-commits --case 1

# Pin the skill runner model and LLM-judge model
arc-skill-eval run ./skills/arc-conventional-commits \
  --model openai-codex/gpt-5.5:medium \
  --judge-model mistral/ministral-8b-latest

# Create a tiny eval-owned Pi config/runtime directory
arc-skill-eval init-runtime ./.arc-skill-eval/pi-agent \
  --provider ollama-cloud \
  --model gpt-oss:20b

# Use an eval-owned Pi config/runtime directory
arc-skill-eval run ./skills/hello-world \
  --agent-dir ./.arc-skill-eval/pi-agent \
  --model ollama-cloud/gpt-oss:20b \
  --judge-model ollama-cloud/gpt-oss:20b

# Generate a static HTML review report and feedback template from run artifacts
arc-skill-eval review ./skills/hello-world/evals-runs/<runId>

# Propose eval improvements from review feedback without writing files
arc-skill-eval improve --from-feedback ./skills/hello-world/evals-runs/<runId>/feedback.json \
  --dry-run --summary

# Retarget output to a different workspace root
arc-skill-eval run . --output-dir ./evals-runs

# Machine-readable JSON
arc-skill-eval run . --json

# Opt into with_skill vs without_skill comparison
arc-skill-eval run . --compare

# Group artifacts under an iteration bucket
arc-skill-eval run . --iteration 1

# Add explicit distractor/conflict skills to the model context
arc-skill-eval run ./skills/arc-conventional-commits \
  --compare \
  --extra-skill ./skills/release-please \
  --iteration conflict-1

# Opt into normal Pi ambient resources such as configured extensions/tools
# while recording the resulting loadout in context-manifest.json
arc-skill-eval run ./skills/arc-conventional-commits \
  --context-mode ambient \
  --iteration ambient-1

Recommended first dogfood run

The companion andysolomon/arc-skills repo ships a dogfood suite for arc-creating-evals, the meta-skill that authors eval suites for other skills. After cloning both repos locally, this is the best end-to-end smoke test:

arc-skill-eval run /path/to/arc-skills/arc-creating-evals \
  --case execution-golden-path-file-skill \
  --model openai-codex/gpt-5.5:medium \
  --judge-model openai-codex/gpt-5.5:medium

For the with-skill / without-skill signal:

arc-skill-eval run /path/to/arc-skills/arc-creating-evals \
  --case execution-golden-path-file-skill \
  --compare \
  --iteration dogfood-1 \
  --model openai-codex/gpt-5.5:medium \
  --judge-model openai-codex/gpt-5.5:medium

A recent dogfood run passed the golden-path case and showed a positive +16.7% with-skill delta after tightening the suite to assert behavior unique to arc-creating-evals.

Create starter evals

Generate a valid starter suite for a skill directory:

arc-skill-eval create ./skills/my-skill

The command reads SKILL.md frontmatter, writes evals/evals.json, and includes three starter cases:

  • trigger-explicit
  • execution-golden-path
  • adjacent-negative

When obvious output artifacts are mentioned in SKILL.md, such as plan.md or report.json, the execution case also gets deterministic file-exists and json-valid assertions. When likely input files are mentioned, such as notes/input.md, requirements.md, prd.md, issue.md, or task.md, the execution case gets seeded fixture inputs under evals/files/starter-inputs/. The adjacent-negative case is domain-aware for common skill types like eval authoring, planning, releases, docs, and auth/webhooks, with a generic fallback. Use --dry-run to print the proposed JSON without writing files, --summary to print a human-readable review of generated cases/assertions, and --force to overwrite an existing evals/evals.json.

Use deterministic create first when the skill has concrete file, JSON, or command-line effects. It is fast, repeatable, CI-friendly, and never spends model tokens. Use create --guided when the hardest part is deciding what to test: conceptual interview skills, planning/review skills, routing skills with subtle adjacent negatives, or skills where success is mostly semantic. Guided mode asks the configured Pi model to design a richer proposal using the bundled skills/arc-creating-evals/SKILL.md procedure, then validates the returned evals.json with the same loader used by run before printing or writing anything. You can pin the designer with --model <provider/model[:thinking]>, use --agent-dir <path> for eval-owned model/auth lookup, or pass --authoring-skill <path> to test a different eval-authoring skill.

Use interactive guided mode to review the proposed suite before it is written:

arc-skill-eval create ./skills/my-skill --guided --interactive

The lightweight prompt flow presents the rationale, cases, fixture inputs, and assertions; lets you include/skip cases and assertions; and lets you edit case prompts, expected output, and judge/regex assertion text. Existing overwrite protections still apply unless --force is supplied.

For example, a conceptual grill-me skill that conducts a relentless interview may not create files at all. Its suite should lean on judge assertions such as "asks direct follow-up questions about assumptions and tradeoffs" plus adjacent negatives that should not trigger the skill, rather than fake file-exists checks. That makes the eval measure the behavior the skill actually promises.

Prefer behavior-focused assertions

Write assertions against observable behavior and artifacts, not incidental wording. Brittle wording checks fail when a correct assistant paraphrases, changes a heading, or omits a phrase the skill never promised.

Prefer:

{ "type": "file-exists", "path": ".releaserc.json" },
{ "type": "regex-match", "pattern": "conventionalcommits", "target": { "file": ".releaserc.json" } },
"The response names semantic-release and explains that it configured release automation for this repository."

Avoid unless the words are truly the product requirement:

"The response says exactly: Phase 1 — detection complete."

Exact wording is appropriate for user-facing contracts such as a required commit message, CLI output, email subject, or safety disclaimer. When wording is required, make it explicit in expected_output and use a deterministic regex-match or exact output assertion so failures explain the missing text directly.

The positional <skill-dir-or-repo> for run is resolved as:

  • a skill directory if it contains evals/evals.json,
  • otherwise a repo whose tree is walked for SKILL.md + evals/evals.json pairs.

Audit skill quality

Run deterministic skill-authoring checks without invoking a model:

arc-skill-eval audit ./skills
arc-skill-eval audit ./skills/my-skill --json
arc-skill-eval audit ./skills --output skill-audit.md

audit reports frontmatter issues, long descriptions, SKILL.md sprawl, missing evals/evals.json, broken local markdown reference links, trigger-heavy descriptions on user-invoked skills, and likely duplicate skill families. It exits successfully by default so it can be used as a report generator; use the finding counts in JSON output if CI needs custom failure thresholds.

Review reports

Turn a run directory into a static review bundle:

arc-skill-eval review ./skills/hello-world/evals-runs/<runId>

This writes review.html and feedback.json into the run directory. Use --output <dir> to write elsewhere and --force to overwrite an existing report. Compare runs are rendered with with_skill and without_skill variants side-by-side.

A practical create-run-review-improve loop looks like this:

arc-skill-eval create ./skills/my-skill --guided --interactive
arc-skill-eval run ./skills/my-skill --case execution-golden-path
arc-skill-eval run ./skills/my-skill --compare --iteration dogfood-1
arc-skill-eval review ./skills/my-skill/evals-runs/iteration-dogfood-1/<runId>
arc-skill-eval improve \
  --from-feedback ./skills/my-skill/evals-runs/iteration-dogfood-1/<runId>/feedback.json \
  --dry-run --summary

Use review.html to inspect assistant output, grading evidence, artifacts, and with/without-skill deltas. Capture human notes in feedback.json; feedback-driven improvement can then turn those notes into a focused plan for changing the skill, tightening assertions, or adding cases.

Improve from feedback

The command reads human notes plus failing assertion summaries and proposes prompt, assertion, fixture, or adjacent-negative changes with rationale. It does not change files unless you pass --apply. Applied changes annotate matching eval cases with validated improvement metadata so the suite remains loadable by run.

Browse runs interactively

Open an interactive terminal run browser (an Ink TUI) over the artifacts under evals-runs/:

arc-skill-eval browse ./skills/arc-conventional-commits   # one skill
arc-skill-eval browse .                                    # whole repo

It renders a lazygit-style four-panel layout — Skills, Cases, Assertions, Runs — with the selected case's prompt, grading evidence, metrics, and with/without-skill comparison in the main pane. It reads the same per-case grading.json / timing.json artifacts that run emits, so no extra setup is needed.

Screenshot of the arc-skill-eval browse terminal UI showing Skills, Cases, Assertions, Runs, and case details

Navigation:

See the Keybindings reference for the full keymap — it's generated from src/tui/keymap.ts, the same source the in-TUI ? overlay renders from, so the two can't drift. Highlights: Tab/1–4 panels, j/k move, →/l/↵ enter the detail pane, [/] cycle case mode, v raw grading.json, / filter, s sort, c pin baseline, r/R run, n new case, ? help, q quit.

Runs and authoring happen in-process, without leaving the TUI:

  • r / R run evals for the selection in a live run console (spinner, per-case progress, pass/fail summary); on completion the affected skill reloads in place and your selection is restored. R adds --compare (with_skill vs without_skill). Esc aborts an in-flight run; ↵ reloads and closes when it's done.
  • o runs with custom flags (--model, --iteration, --extra-skill…). This is the one path that still uses a child process — the binary is arc-skill-eval (must be on PATH); override with ARC_SKILL_EVAL_BIN.
  • n scaffolds a new eval case into evals/evals.json; f records a feedback.json note for the selected case (consumed by improve).

Display options:

  • --no-baseline hides the without_skill comparison rows in the detail pane (handy when you only ran the skill variant).
  • The TUI is capability-aware: truecolor hex degrades to 16-color ANSI on low-color terminals, and block/box glyphs (bars, status ticks, accent bar) fall back to ASCII off a UTF-8 locale. Force the fallbacks with NO_COLOR / FORCE_COLOR=0 (no color) or ARC_TUI_ASCII=1 (ASCII glyphs).

Model options:

  • --model <provider/model[:thinking]> pins the skill runner model instead of using Pi's configured default. Example: openai-codex/gpt-5.5:medium.
  • --judge-model <provider/model[:thinking]> pins the model used for LLM-judged string assertions. Deterministic assertions do not use the judge.
  • --agent-dir <path> points Pi settings, model registry, and auth lookup at an eval-owned agent directory instead of the normal ~/.pi/agent directory.
  • When no model flags are supplied, arc-skill-eval inherits Pi's default provider/model/thinking level from the effective Pi agent settings.

Export results to Laminar Evaluations (optional)

run --laminar additionally reports the run to Laminar's Evaluations view: one evaluation per run variant (with_skill / without_skill), one scored datapoint per case, grouped by skill name so variants can be compared side by side. It is off by default and entirely optional — local evals-runs/ artifacts remain the canonical record, and an export failure never fails the run.

LMNR_PROJECT_API_KEY=lmnr_... arc-skill-eval run . --laminar

The run summary prints a direct dashboard link per evaluation. Each datapoint carries numeric scores (pass_rate, passed, failed, total_tokens, cost_usd, duration_ms, tool_calls) and an output with the grading summary, per-assertion verdicts (assertion text, pass/fail, short evidence quote), and local artifact paths.

  • LMNR_PROJECT_API_KEY is required when --laminar is set; the command fails fast (before any case runs) and names the key if it is missing. LMNR_BASE_URL is optional; LMNR_PROJECT_NAME optionally overrides the evaluation group name (default: the skill name).
  • The Laminar Node SDK (@lmnr-ai/lmnr) is an optional dependency, dynamically imported only when the flag is enabled — installs that never use Laminar don't pull it in. If it's missing when enabled, the run reports a clear error naming the package.
  • Exports carry grading verdicts, metrics, and artifact paths only (never full assistant text, prompts, or file contents). The benchmark.json delta remains local.

See docs/concepts/artifacts → External observability for the full local-artifact → Laminar evaluation mapping.

Eval-owned Pi runtime

Use --agent-dir when you want reproducible team or CI runs without depending on personal Pi defaults:

arc-skill-eval run ./skills/hello-world \
  --agent-dir ./.arc-skill-eval/pi-agent \
  --model ollama-cloud/gpt-oss:20b \
  --judge-model ollama-cloud/gpt-oss:20b

Create one with:

arc-skill-eval init-runtime ./.arc-skill-eval/pi-agent \
  --provider ollama-cloud \
  --model gpt-oss:20b

Use --force to intentionally overwrite existing runtime files.

A minimal eval-owned runtime contains just:

.arc-skill-eval/pi-agent/
├── models.json
└── settings.json

The runner and default LLM judge both use this directory for Pi models.json, settings.json, and auth.json lookup when --agent-dir is supplied. run preflights this directory before executing cases and reports missing models.json, settings.json, provider/model entries, or required API-key environment variables once with an init-runtime remediation. Secrets should still be referenced by environment variable name, for example "apiKey": "OLLAMA_API_KEY", rather than committed as literal values.

Ollama / low-cost cloud and local runs

arc-skill-eval inherits model support from Pi. Ollama Cloud is a useful low-cost provider lane for smoke tests. A verified working example is:

arc-skill-eval run ./skills/hello-world \
  --model ollama-cloud/gpt-oss:20b \
  --judge-model ollama-cloud/gpt-oss:20b

A recent run with ollama-cloud/gpt-oss:20b passed 2/3 hello-world cases. The failed case was model behavior on an ambiguous prompt, not provider failure: the model asked which name to use instead of defaulting to Hello, world!.

Pi can also be configured through Ollama's integration for local or proxied cloud models:

# Let Ollama install/configure Pi and launch an interactive session
ollama launch pi

# Configure Pi for Ollama without launching
ollama launch pi --config

# Example cloud model launch through Ollama
ollama launch pi --model qwen3.5:cloud

After Pi lists Ollama models, use the same provider/model pinning flags:

arc-skill-eval run ./skills/hello-world \
  --model ollama/qwen3.5:cloud \
  --judge-model ollama/qwen3.5:cloud

For direct Ollama Cloud access, set OLLAMA_API_KEY and add an ollama-cloud provider to Pi's models.json:

{
  "providers": {
    "ollama-cloud": {
      "baseUrl": "https://ollama.com/v1",
      "api": "openai-completions",
      "apiKey": "OLLAMA_API_KEY",
      "models": [
        { "id": "gpt-oss:20b" },
        { "id": "ministral-3:3b" },
        { "id": "gemma3:4b" }
      ]
    }
  }
}

For local Ollama setup, add an Ollama-compatible provider to ~/.pi/agent/models.json using http://localhost:11434/v1 and set defaultProvider / defaultModel in ~/.pi/agent/settings.json. Local models do not require OLLAMA_API_KEY.

For the runtime roadmap — a tiny eval-owned Pi config vs a future custom agent — see docs/agent-runtime-strategy.md.

Context options:

  • --extra-skill <path> can be repeated to add explicit skill directories or SKILL.md files as distractor/conflict context. In --compare, with_skill receives the target + extras, while without_skill receives extras only.
  • --context-mode isolated is the default: no ambient Pi skills, extensions, prompt templates, themes, or context files are loaded.
  • --context-mode ambient opts into normal Pi ambient resources so extension tools/MCP-like tools and other configured resources can enter the context. The resolved loadout is recorded in context-manifest.json.
  • --sandbox none|just-bash selects the execution isolation for every selected case, overriding each case's own sandbox field. none (default) uses the temp-workspace runner; just-bash routes the agent's bash tool through an in-process virtual shell (filesystem rooted at the case workspace) so commands run without the host shell and never touch the repo tree. Generated files are still captured under outputs/. npm/npx/git resolve to deterministic mocks (no-op success by default, configurable per case via sandboxMocks).

Exit code: 0 when every case has no failing assertions, 1 otherwise.

Output layout

For each default single-variant run:

<skillDir>/evals-runs/<runId>/
├── eval-<case-id>/
│   ├── assistant.md          # final assistant response text
│   ├── outputs/              # files produced by the run
│   ├── timing.json           # duration, model, thinking, token/cost/context metrics
│   ├── grading.json          # per-assertion passed + evidence
│   ├── trace.json            # normalized runtime trace + raw telemetry refs
│   ├── tool-summary.json     # tool calls, errors, skill reads, external/MCP activity
│   └── context-manifest.json # skills/tools/context exposed to the model

Use --iteration <name> to group artifacts under <skillDir>/evals-runs/iteration-<name>/<runId>/; for example --iteration 1 writes to iteration-1/<runId>/.

With --compare, each case writes isolated variant artifacts and the skill run root includes benchmark.json:

<skillDir>/evals-runs/<runId>/
├── benchmark.json            # with_skill vs without_skill aggregate
├── eval-<case-id>/
│   ├── with_skill/
│   │   ├── assistant.md
│   │   ├── outputs/
│   │   ├── timing.json
│   │   ├── grading.json
│   │   ├── trace.json
│   │   ├── tool-summary.json
│   │   └── context-manifest.json
│   └── without_skill/
│       ├── assistant.md
│       ├── outputs/
│       ├── timing.json
│       ├── grading.json
│       ├── trace.json
│       ├── tool-summary.json
│       └── context-manifest.json

timing.json includes runner observability:

{
  "total_tokens": 12345,
  "duration_ms": 50123,
  "model": { "provider": "anthropic", "id": "claude-opus-4-5", "thinking": "medium" },
  "thinking_level": "medium",
  "token_usage": {
    "input_tokens": 10000,
    "output_tokens": 2000,
    "cache_read_tokens": 300,
    "cache_write_tokens": 45,
    "total_tokens": 12345
  },
  "estimated_cost_usd": 0.1234,
  "context_window_tokens": 200000,
  "context_window_used_percent": 6.2
}

tool-summary.json highlights behavior-level observability:

{
  "tool_call_count": 8,
  "tool_error_count": 0,
  "tool_calls_by_name": { "read": 2, "bash": 3, "write": 2, "edit": 1 },
  "skill_read_count": 1,
  "skill_reads_by_name": { "arc-conventional-commits": 1 },
  "external_call_count": 0,
  "mcp_tool_call_count": 0
}

context-manifest.json records the run loadout so skill/tool conflicts can be diagnosed:

{
  "runtime": "pi",
  "mode": "isolated",
  "attached_skills": [{ "name": "arc-conventional-commits", "path": ".../SKILL.md", "role": "target" }],
  "available_tools": [{ "name": "bash", "source": "builtin" }],
  "active_tools": ["read", "bash", "edit", "write"],
  "mcp_tools": [],
  "mcp_servers": [],
  "ambient": { "extensions": false, "skills": false, "prompt_templates": false, "themes": false, "context_files": false }
}

grading.json per the Anthropic format:

{
  "case_id": "1",
  "assertion_results": [
    { "text": "file-exists: .releaserc.json", "passed": true, "evidence": "Found .releaserc.json (182 bytes)", "assertion": { "type": "file-exists", "path": ".releaserc.json" } },
    { "text": "The response summarizes the semantic-release plugins it installed.", "passed": true, "evidence": "\"installs @semantic-release/commit-analyzer + release-notes-generator\"", "assertion": "The response summarizes the semantic-release plugins it installed." }
  ],
  "judge_model": { "provider": "mistral", "id": "ministral-8b-latest" },
  "summary": { "passed": 2, "failed": 0, "total": 2, "pass_rate": 1.0 }
}

judge_model records the LLM-judge that graded the prose assertions; it's omitted for cases with only deterministic checks. browse surfaces it per case. Precedence: the --judge-model selection, else the model that ran the case (which is known to be authenticated), else { "provider": "mistral", "id": "ministral-8b-latest" } as a last resort. Note that when the judge defaults to the runner's model, the model grades its own output — pin --judge-model to a different model when that bias matters.

Authoring an eval suite for a skill

Use the bundled arc-creating-evals skill in skills/arc-creating-evals/. It interviews you across Anthropic's four success dimensions (outcome, process, style, efficiency) and emits evals/evals.json + fixtures. Install the skill into your agent's skills directory (.claude/skills/ or the equivalent for your tool) — see skills/README.md for the recipe.

Docs

  • docs/skill-eval-authoring-debrief.md — detailed research debrief and playbook for creating evals for skills, including the arc-skills mastery roadmap.
  • docs/agent-runtime-strategy.md — runtime strategy for Pi-backed evals, tiny eval-owned Pi config, Ollama Cloud, and a possible future custom agent.
  • docs/skill-creator-parity-plan.md — user stories and implementation plan for Claude skill-creator parity features.
  • docs/skill-creator-parity-progress.txt — checkbox tracker for the skill-creator parity roadmap.
  • docs/create-dogfood-report.md — findings from dry-running create against real arc-skills.
  • docs/evals-json-pivot.md — direction, milestone log, and what stays vs what was deprecated.
  • docs/domain-model.md — runtime + grading entities.

Shipped since the MVP

The slim MVP of the pivot to the Anthropic format has since grown both of its planned follow-ups:

  • Cross-iteration comparison — pin a baseline run with c in browse to diff iterations against it, and R runs --compare (with_skill vs without_skill) in place.
  • Human-review feedback.json — written by review (and by f in browse), consumed by improve --from-feedback.

Remaining ideas live in docs/evals-json-pivot.md and ROADMAP.md.

Inspiration & credits

arc-skill-eval exists because three pieces of writing made it clear what to build, in what shape, and with what philosophy. Each one shaped a different layer:

  • The eval methodology — every workflow choice in the grader and the suite-growth advice in the docs — comes from OpenAI's Testing Agent Skills Systematically with Evals by Dominik Kundel and Gabriel Chua (January 22, 2026). The framing of an eval as "a prompt → a captured run (trace + artifacts) → a small set of checks → a score you can compare over time", the layered-grading recipe (fast deterministic checks first, then model-assisted rubric), the multi-category success metrics (outcome / process / style / efficiency), and the guidance that "a small set of 10–20 prompts is enough to surface regressions" — these are OpenAI's, transposed onto Anthropic's published format.
  • The eval format — the on-disk evals/evals.json shape, the per-case grading.json, the aggregate benchmark.json, and the with_skill / without_skill comparison — comes from Anthropic's documented skill-eval methodology. The framework consumes Anthropic's format so a skill author can take their evals.json to any compatible runner.
  • The runtime philosophy — the bias toward a small, legible runtime that's not afraid to call itself a loop — owes a debt to two posts that demystified the agentic harness:
    • Thorsten Ball's How to Build an Agent (Ampcode, April 15, 2025) — "It's an LLM, a loop, and enough tokens" — and the demonstration that a useful code-editing agent fits in a few hundred lines.
    • Mihail Eric's The Emperor Has No Clothes: How to Code Claude Code in 200 Lines of Code (January 2026), which makes the same point at the level of agent harnesses: the core is a tool registry, an inner loop, and a parser. Production complexity is engineering, not architecture.

If you read only one of those before authoring an eval, read OpenAI's. If you read only one before extending the framework, read either of the harness pieces.