npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@tangle-network/agent-bench

v0.13.3

Published

Benchmark adapters and execution for agent-runtime across coding, tool-use, RAG, memory, browser, and terminal tasks.

Readme

agent-runtime-bench

Published as @tangle-network/agent-bench, with independent CI and release checks for its TypeScript and Python surfaces.

Read bench/HARNESS.md FIRST. It is the one maintained map: the commands, the rollout → corpus → selector → CI → gate data flow, the canonical-suite table, the wired/needs-creds/scaffolded matrix, and the gate one-liners, kept verified against source.

Release

This package publishes from its own tag. A root v* tag publishes @tangle-network/agent-runtime only.

# after bench/package.json holds the new version on the tip of main
git tag -a agent-bench-v<version> -m "agent-bench <version>"
git push origin agent-bench-v<version>

Bump bench/package.json and cut the tag in the same release. The manifest resolves the first-party cohort through the workspace catalog, so a catalog move changes what this package publishes. publish.yml skips a version the registry already holds, so a bump without a tag leaves the registry serving the previous cohort to every consumer.

Confirm the release with npm view @tangle-network/agent-bench version.

Use

pnpm add -D @tangle-network/agent-bench
import { resolveAdapter } from '@tangle-network/agent-bench'

const crag = resolveAdapter('crag')

SWE-bench judge setup (the one block not in HARNESS.md)

python3 -m venv .venv && .venv/bin/pip install swebench   # SWE-bench harness
pnpm install                                              # tsx + link parent
# Docker daemon must be running (judges build/run per-instance images)

The judge needs only Docker; workers need a model key (Tangle router TANGLE_API_KEY, or a direct provider).

Live optimizer scripts require explicit token prices so cost records cannot be guessed. Set REFLECT_INPUT_USD_PER_MILLION, REFLECT_CACHED_INPUT_USD_PER_MILLION, REFLECT_CACHE_WRITE_USD_PER_MILLION, and REFLECT_OUTPUT_USD_PER_MILLION. Missing prices fail before an optimizer model call.

Retain every official per-test log and report before the temporary evaluator directory is removed:

const adapter = createSweBenchAdapter({
  captureEvaluatorArtifacts: ({ taskId, attemptSequence }) => ({
    destination: path.join(runDirectory, taskId, String(attemptSequence)),
  }),
})
const score = await adapter.judge(task, patch)
console.log(score.judgeArtifacts?.manifestPath)

Each destination contains the untouched evaluator tree under evaluator/, raw stdout.bin and stderr.bin under process/, and receipt.json with per-file SHA-256 values plus a whole-tree SHA-256. The destination must be unique and absent; an existing path fails loud instead of overwriting evidence. Failed evaluators throw StagedJudgeError with the same judgeArtifacts receipt after retaining partial logs.

Official-data adapters: tau2/tau3, DABStep, FinResearchBench

Fixture mode (TAU2_FIXTURES=1, TAU3_FIXTURES=1, DABSTEP_FIXTURES=1, FINRESEARCHBENCH_FIXTURES=1) proves adapter plumbing only. A fixture result is never a benchmark score. A benchmark score requires the official data below plus the benchmark's own judge, and nothing else counts.

| Adapter | Official data | Judge | |---|---|---| | tau2-bench / tau3-banking | TAU2_BENCH_DIR / TAU3_BENCH_DIR = clean git checkout of sierra-research/tau2-bench; domain via TAU2_DOMAIN / TAU3_DOMAIN | tau2's own reward recomputation over a tau results.json trajectory | | dabstep | DABSTEP_DIR = checkout of EnvCommons/DABStep (grade.py, splits/, files/); released dataset.csv beside it or via DABSTEP_DATASET_CSV | official grade.py, deterministic, no LLM | | finresearchbench | FINRESEARCHBENCH_DATA_FILE = JSON/JSONL export whose rows carry judge_system_prompt and judge_prompt_template | official logic-tree model judge over the Tangle router (TANGLE_API_KEY) |

tau2/tau3. Run uv sync --extra knowledge inside the checkout and set AGENT_BENCH_PYTHON=<checkout>/.venv/bin/python3 so the loader and judge run in the interpreter that holds the tau2 distribution. Every official load stamps upstreamCommit (checkout HEAD) and upstreamVersion (installed tau2 distribution) into task metadata, and fails loud on a dirty or non-git checkout or a missing distribution. The judge re-resolves the pin, refuses a checkout or interpreter that moved after load, and writes both values into its score detail. A trajectory score needs a tau results.json produced by the paid tau simulation (agent plus user simulator); the adapter only recomputes the official reward from that artifact.

DABStep. The git checkout does not ship dataset.csv. The row file (task_id,question,guidelines,all_golds_by_task) is distributed with the OpenReward DABStep environment, which mounts it at /orwd_data/dataset.csv. The public adyen/DABstep release carries the task rows without golds (data/tasks/all.jsonl; answers withheld for the leaderboard) and a 10-task dev split with reference answers (data/tasks/dev.jsonl). Point DABSTEP_DATASET_CSV at an absolute row-file path when the checkout and the rows live apart.

FinResearchBench. The judge is a model call, so its score records the exact judge turn usage on BenchScore.judgeUsage (tokens, cost, model). tokensKnown: false and usdKnown: false survive verbatim; an unknown judge cost never reads as zero spend. FINRESEARCHBENCH_JUDGE_MODEL (or JUDGE_MODEL) selects the judge; FINRESEARCHBENCH_PASS_THRESHOLD moves the resolve bar from 0.8.

Pier custom candidates

The package executes a branded PreparedAgentCandidateExecution from @tangle-network/agent-runtime through one atomic API and ships pier_agents.tangle_candidate:TangleCandidateAgent as its thin Pier transport. The executor recreates every input from runtime-verified file bytes and reveals model credentials only inside the claimed execution callback. Pier owns the task container and verifier; protected model usage and traces stay in @tangle-network/agent-eval and are finalized by the shared runtime. FilePierCandidateTrialController atomically reserves a unique Pier job, then persists the supervisor PID, process-session identity, and that job's exact Docker projects so a fresh evaluator process can stop and remove an abandoned trial. Run PIER_REPO=/path/to/pier pnpm verify:pier for the zero-model failure/pass and fresh-process recovery proof, and see HARNESS.md for the exact invocation and failure contract. From an installed npm package, expose the shipped Python module with export PYTHONPATH="$(npm root)/@tangle-network/agent-bench${PYTHONPATH:+:$PYTHONPATH}" before invoking Pier.

Control execution within a managed shot

Supply execute to control whole sandbox prompts while Bench owns setup, extraction, grading, and cleanup. Your policy can inspect permitted working checks through the existing sandbox handle.

import { runBenchmarks, type BenchExecution } from '@tangle-network/agent-bench'
import { executionOptions, chooseNextPrompt } from './policy.js'

const execute: BenchExecution = async ({ run, prompt, signal }) => {
  const first = await run.start(prompt)
  const correction = await chooseNextPrompt({ first, box: run.box, sessionId: run.sessionId, signal })
  if (correction !== undefined) await run.resume(correction)
}

const report = await runBenchmarks({ ...executionOptions, execute })

The consumer supplies and identifies chooseNextPrompt; Bench provides no learning policy. perTask[].prompts retains each invocation and its observed worker cost, including partial failures. See the execution contract for accounting, retry composition, and lifecycle limits.