npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@nimblehq/claude-sdk

v1.7.0

Published

Anthropic Claude Agent SDK automations (spec, tasks, implement, review, etc.)

Readme

Claude SDK automations

Anthropic-backed automations: spec, task breakdown, implement, create-pr, review-pr, test gap, changelog, release gate, and CLAUDE.md generation.

Consumer docs: GitHub Actions · npm / local


Stages

| Folder | Stage | Purpose | |---|---|---| | spec-generator/ | spec_generator | Request → cross-platform PRD | | task-breakdown/ | task_breakdown | Spec → sprint-ready tasks | | implement/ | implement | Branch, code, commit, push | | create-pr/ | create_pr_draft + create_pr | Draft body, open PR as draft, ensure labels, request CODEOWNERS reviewers, assign latest milestone | | review-pr/ | review_pr | Inline review; structured verdict + deletions recommended; always COMMENT | | test-gap-detector/ | test_gap_detector | Missing-test coverage notes on PR diffs | | changelog-generator/ | changelog | Changelog from commits since last release tag | | release-gate/ | release_gate | Pre-release safety check | | claude-md-generator/ | claude_md_generator | Generate / refresh CLAUDE.md | | pipeline/ | — | Matrix builder + summary for full-pipeline workflow |


Pull requests

create_pr opens as a draft; only a human marks it ready. Also best-effort: ensure type : * labels, CODEOWNERS reviewers, latest open milestone.

implement runs one cleanup + verification round before create_pr — see pr-checklist.md.


Models per stage

Each stage picks a model family, not a pinned version string:

| Stage | Family | |---|---| | triage, spec_generator, task_breakdown, implement, fix_pr_comments, claude_md_generator | Sonnet | | create_issue_draft, create_issue, create_pr_draft, create_pr, test_gap_detector, changelog | Haiku | | review_pr, release_gate | Opus |

Defined in lib/agent-utils.ts (MODEL_FAMILY_PER_STAGE). modelForStage() resolves each family to the actual latest model ID via the Anthropic Models API (lib/model-resolver.ts, client.models.list()) — new model releases are picked up automatically, no code change or redeploy needed. Resolution is cached per process (one Models API call per stage run, not per message); if the Models API is unreachable, each family falls back to a last-known-good pinned ID in lib/model-resolver.ts. Fable is never selected — deliberately excluded as too expensive for this pipeline's per-stage budgets (see Cost & timeout gates). Override by passing model to a stage input.


Knowledge layer

Nimble conventions injected into every prompt by loadKnowledge(projectRoot). SDK defaults in knowledge/shared/; projects override per-file in <repo>/.claude/knowledge/. Hard cap: 32 KB total.

| File | Purpose | |---|---| | git-conventions.md | Branch / commit / PR title format | | pr-checklist.md | Definition of done for drafter + reviewer | | github-conventions.md | Issue/PR title, label, milestone, assignee conventions | | coding-standards.md | Implementation conventions, tests, error handling | | security-guidelines.md | Secrets, auth, deps, logging |


Skills

Loaded from skills/ via combineSkills() in config/skills.ts.

| Skill | Stage | |---|---| | code-review-and-quality | review_pr | | security-and-hardening | review_pr, release_gate | | thermo-nuclear-code-quality-review | review_pr | | spec-driven-development | spec_generator | | api-and-interface-design | spec_generator | | planning-and-task-breakdown | task_breakdown | | incremental-implementation | implement (feature, chore) | | test-driven-development | implement (bug, feature) | | debugging-and-error-recovery | implement (bug) | | shipping-and-launch | release_gate | | documentation-and-adrs | changelog |


Shared modules

| Module | Purpose | |---|---| | lib/agent-utils.ts | MODEL_PER_STAGE, modelForStage(), withModelFallback(), stage timeouts, lifecycle logging | | lib/knowledge.ts | Two-layer loader (SDK shared + project overrides); systemWithKnowledge() for agent stages | | lib/state.ts | .claude-pipeline-state.json load/save; includes traceId; loadState() validates all required fields and throws on a corrupt file | | lib/git-env.ts | resolveGithubRepo, resolveBaseBranch (auto-detect from git) | | config/client.ts | Anthropic SDK client; re-exports MODEL_PER_STAGE | | config/skills.ts | Loads skills from skills/; skillsForWorkType per-workType composition | | config/stacks.ts | Stack hints for implement + review_pr | | config/cost-tracker.ts | Per-run cost logging to .observability/costs.jsonl | | config/cached-prompt.ts | cachedSystem() — pure prompt-cache block helper (no API client) | | config/create-cached-message.ts | createCachedMessage() — Messages API call with cache + usage logging | | config/batch.ts | runBatchMessage() — Batch API wrapper with polling, fallback, and env-var tuning |


Environment variables

| Variable | Required | Description | |---|---|---| | ANTHROPIC_API_KEY | yes | Anthropic API key | | GH_BOT_TOKEN | recommended | PAT with repo + workflow scopes for git push and gh CLI. Falls back to github.token when unset. | | LINEAR_API_KEY | no | Routes spec publishing to Linear | | CLAUDE_SDK_DEBUG | no | Set to 1 to dump raw Anthropic SDK events to stderr | | CLAUDE_MAX_USD_PER_STAGE | no | Global USD cap per stage. Default: see Cost & timeout gates below | | CLAUDE_MAX_USD_<STAGE> | no | Per-stage USD override (e.g. CLAUDE_MAX_USD_IMPLEMENT=15). Wins over the global cap | | DISABLE_BATCH_API | no | Set to true to fall back to synchronous messages.create() for changelog and test_gap_detector | | BATCH_POLL_INTERVAL_MS | no | Milliseconds between Batch API status checks (default 10000) | | BATCH_MAX_ATTEMPTS | no | Max Batch API poll iterations before timing out (default 12, ~120 s) |


Cost & timeout gates

Two hard gates keep a runaway stage from burning unbounded API spend:

  • Wall-clock timeout per stage — implement 30 min, review_pr / release_gate 15 min, drafts 10 min. Passed to runWithTimeout() at each call site.
  • USD budget — every result message from the agent stream is checked against maxBudgetForStage(stage); the stage throws as soon as total_cost_usd exceeds the cap.

| Stage | Default cap | |---|---| | implement | $10 | | review_pr, release_gate | $5 | | all other stages | $2 |

Override with CLAUDE_MAX_USD_<STAGE> (per stage) or CLAUDE_MAX_USD_PER_STAGE (global). The per-stage value wins.

When $GITHUB_STEP_SUMMARY is set, each stage appends a Stage | Duration | Cost | Runs | Status row so the spend is visible on the Actions run page.


Tool admission control (makeToolPolicy)

Every query() call site passes a stage-bound hook built by makeToolPolicy({ stage, traceId }) from hooks/safety.ts. Two layers run on each tool call:

| Layer | Rule | |---|---| | Per-stage allowlist | A stage may only call the tools in TOOL_ALLOWLIST[stage], enforced even if allowedTools is widened by mistake | | Bash denylist | Destructive shell commands are blocked before they reach the runner |

TOOL_ALLOWLIST: implement → Read, Write, Edit, Bash, Glob, Grep; review_pr → Read, Glob, Grep; create_pr_draft / create_pr → none (text only). Denylist patterns are sourced from automations/shared/safety-denylist.ts — the same list Cursor SDK's hooks consume. Blocks include rm -rf ~/.ssh, cat .env, curl … | sh, gh secret set, git filter-branch, --no-verify, force-push, and fork bombs.

Every decision (allow and deny) is appended to .observability/tool-calls.jsonl with the stage, traceId, tool, and reason — a forensic record of what each agent asked to do. Override the path with TOOL_AUDIT_LOG_PATH.

The local copy at safety-denylist.ts is regenerated from the canonical file via npm run prebuild. Do not edit the local copy.

MCP (optional): githubMcpServerConfig() in lib/mcp-github-server.ts (synced from automations/shared/mcp-github-server.ts) returns config for the official github-mcp-server — the same image the repo's daily-repo-status gh-aw workflow already runs. Returns undefined unless MCP_GITHUB_SERVER=1 and a GitHub token are set, so it's opt-in and not wired into any stage's query() call by default. This is the structured-tool-call alternative to reaching for Bash + gh in stages that need GitHub API access.


Input sanitization

sanitize.ts (autogenerated from automations/shared/sanitize.ts via npm run prebuild) guards user-controlled text before it reaches any LLM prompt.

sanitizeUntrustedInput({ text, source }) — for free-form user content:

  1. Strips <tool>, <function_calls>, <invoke>, <result>, <param> XML markup (tag content preserved).
  2. Checks 16 denylist patterns — instruction overrides, role-change attacks, jailbreak keywords, system-tag injection. Returns [TRUNCATED: <source>] and logs [SECURITY] on match.
  3. Truncates at 4 000 chars and wraps clean text in <UNTRUSTED_USER_INPUT source="…"> boundary tags.

sanitizePipelineHint(value, source) — for short workflow metadata:

  • Strips newlines, caps at 200 chars, applies the same denylist. Returns empty string on a match.

Currently wired into: implement (WORK_ITEM_TITLE, each ACCEPTANCE_CRITERIA), review_pr (PR_TITLE, PR_DESCRIPTION), spec_generator (SPEC_REQUEST, SPEC_CONTEXT).

Rule of Two. The pipeline simultaneously processes untrusted free-text (request, issue bodies, PR comments) and holds write-capable secrets and PR-creation ability — the combination the industry calls out as the risky one for agentic workflows. sanitizeUntrustedInput/sanitizePipelineHint above are the mitigation in place today; treat any new stage that reads untrusted text as needing the same sanitization before it reaches a prompt.


Scripts

| Script | Purpose | |---|---| | scripts/dev/pipeline.ts | Full pipeline CLI: spec → breakdown → implement → create-pr | | scripts/dev/test-pipeline.ts | Dry run (stages 1–2 only) | | scripts/dev/smoke-test.ts | Verifies ANTHROPIC_API_KEY, SDK wiring, and prompt-cache read hits | | scripts/dev/test-implement.ts | Local implement test against TEST_REPO_PATH | | scripts/dev/test-review-pr.ts | Local review_pr smoke test | | scripts/dev/test-changelog-generator.ts | Local changelog test | | scripts/dev/test-spec-generator.ts | Local spec generator test | | scripts/dev/test-task-breakdown.ts | Local task breakdown test | | scripts/dev/test-release-gate.ts | Local release gate test | | scripts/dev/test-test-gap-detector.ts | Local test-gap detector test | | scripts/dev/test-claude-md-generator.ts | Local CLAUDE.md generator test | | npx claude-sdk-cost-report | Summary from .observability/costs.jsonl |


Observability

Cost log: .observability/costs.jsonl (gitignored). Every stage run appends one JSONL record via config/cost-tracker.ts (withCostTracking()). Fields: automation, runId, traceId, tokens, cost estimate, cache write/read tokens, duration, success/failure.

traceId is a UUID generated by pipeline/build-implementation-matrix.ts and emitted as trace_id to GITHUB_OUTPUT. The workflow passes it to each stage job as TRACE_ID; action files read it and thread it to withCostTracking() so all costs for one pipeline invocation share a common ID.

Summary: npx claude-sdk-cost-report

The report includes per-automation cache hit rates, a write-to-read ROI line, and a break-even annotation. If cacheCreationTokens is 0 across all runs, the report prints a setup reminder — prompt caching is not active.

Every agent-running CI job uploads .observability/costs.jsonl as an observability-<runId>-<stage> artifact (30-day retention, if: always()). Download from the Actions run page to debug cost spikes or timeouts post-mortem.

Prompt caching

Prompt caching is enabled via ANTHROPIC_API_KEY only — no beta headers or extra secrets. See Anthropic prompt caching docs.

Messages API stages mark the stable system block with cache_control: { type: 'ephemeral' }:

  • Non-batch stages (spec_generator, task_breakdown, release_gate, claude_md_generator) call createCachedMessage() from config/create-cached-message.ts — it also picks the stage model via modelForStage() and logs usage.
  • Batch stages (changelog, test_gap_detector) wrap the system with cachedSystem() from config/cached-prompt.ts (kept client-free so it imports without ANTHROPIC_API_KEY).

Cache usage is recorded in .observability/costs.jsonl via recordMessageUsage().

Minimum cacheable length: caching only activates above the model's threshold — 1,024 tokens for Sonnet 4.6 / Opus 4.8, but 4,096 for Haiku 4.5. The Haiku stages (changelog, test_gap_detector) sit near or below that, so caching there is marginal-to-no-op; the reliable savings are on the Sonnet/Opus stages.

Agent SDK stages (implement, review_pr, create_pr_draft) rely on the Agent SDK's internal cache markers. Stable content (skills + project knowledge from systemWithKnowledge()) lives in systemPrompt; per-task details stay in prompt so matrix jobs within a pipeline run share the same cached prefix.

Verify locally:

npx tsx scripts/dev/smoke-test.ts   # two identical calls; 2nd should show cache_read > 0
npx claude-sdk-cost-report          # per-stage cache write/read tokens

Break-even at Sonnet 4.6 pricing ($3.00/MTok input, $3.75/MTok write, $0.30/MTok read): >12.5 reads per write. Stages triggered multiple times per day will exceed this threshold within hours of a cache cold start.


Publishing (maintainers)

Driven by sdk-publish.yml, triggered on a published GitHub Release with a claude-sdk/vX.Y.Z tag.

One-time setup

  • secrets.NPM_TOKEN — npm automation token with publish access to the @nimblehq scope.
  • The @nimblehq scope must exist on npm and the publishing user must have developer/owner access.

Procedure

  1. Bump version in package.json and add a matching entry to CHANGELOG.md. Merge to main.
  2. Run Create GitHub Release → select claude-sdk. Version is read from package.json — no input needed.

The workflow reads the version from automations/claude-sdk/package.json, creates the claude-sdk/vX.Y.Z tag, publishes the release, and triggers sdk-publish.yml to build and push to npm and GitHub Packages. sdk-publish.yml re-checks that the tag matches package.json before publishing.


Diagnostics

| File | Contents | |---|---| | .claude-pipeline-state.json | Stage handoff state | | .claude-pipeline-runs.jsonl | Stage and run lifecycle events | | .observability/costs.jsonl | Per-run cost log (also uploaded as observability artifact) | | .observability/tool-calls.jsonl | Per-tool-call audit log: stage, traceId, tool, allow/deny decision (also uploaded) | | .claude-pipeline-review-raw-<ts>.txt | Raw review_pr output when JSON parsing fails |