@nimblehq/cursor-sdk
v1.8.0
Published
Cursor SDK automations (triage, implement, PR tools)
Readme
Cursor SDK automations
Pipeline driven by CURSOR_API_KEY, running against a checked-out target repo. Independent of automations/claude-sdk/.
Consumer docs: GitHub Actions guide · npm / local guide
Execution model
Each workflow job:
- Checks out this repo (the platform).
- Checks out the target repo at
target-repo/viacursor-bootstrap. - Sets git committer to
Nimble Cursor Bot <[email protected]>. - Seeds a per-run support bundle into the target repo:
| Bundle file | Tracked by project? |
|---|---|
| .cursor/skills/* | No (gitignored per-run) |
| .cursor/hooks.json + .cursor/hooks/*.js | No (gitignored per-run) |
| .cursor/knowledge/<file>.md | Yes — seeded from knowledge/shared/ only when missing; project commits overrides |
Models per stage
lib/model-catalog.ts's resolveLatestModel() resolves each family's current plain (non-thinking) model from Cursor.models.list() at runtime, falling back to a pinned id if the call fails or times out (20s).
| Stage | Model family | Resolves to (as of writing) |
|---|---|---|
| triage, fix-pr-comments | sonnet | claude-sonnet-5 |
| implement | grok | grok-4.5 |
| review-pr | opus | claude-opus-5 |
| create_issue_draft, create_issue, create_spec, create_pr_draft, create_pr | default (Cursor's own "Auto" selection) | — |
Pull requests
create_pr opens as a draft; only a human marks it ready. Also best-effort: ensure type : * labels, CODEOWNERS reviewers, latest open milestone.
implement runs one cleanup + verification round before create_pr — see pr-checklist.md.
Environment variables
| Variable | Required | Description |
|---|---|---|
| CURSOR_API_KEY | yes | Cursor API key |
| GH_BOT_TOKEN | recommended | PAT for git push + gh CLI. Falls back to github.token; github.token cannot push under .github/workflows/ |
| SHORTCUT_API_TOKEN | no | Routes issues to Shortcut (priority over Linear) |
| LINEAR_API_KEY | no | Routes issues to Linear |
| LINEAR_TEAM_ID | no | Linear team ID; required when LINEAR_API_KEY is set |
| CURSOR_SDK_DEBUG | no | Set to 1 to dump raw Cursor SDK stream events to stderr |
| CURSOR_MAX_RUNS_PER_STAGE | no | Max agent.send() runs per stage (default 20). See Run gates below |
| CURSOR_MAX_AGENTS_PER_STAGE | no | Max Agent.create() calls per stage (default 3) |
| RETRY_MAX_ATTEMPTS | no | Transient-error retry attempts per agent lifecycle (default 2). Retries on 429, 5xx, ECONNRESET, socket hang up, fetch failed. Terminal errors (401, 403, 404, budget cap) never retry |
| CURSOR_ALLOW_CLOUD_UNGATED | when cloud | Required to be 1 when runtime: "cloud" is used. Cloud sandboxes bypass .cursor/hooks/*.js, so block-destructive / approve-pr / log-edit gates do NOT apply |
Run gates
Cursor SDK does not surface per-message tokens or USD cost, so cost containment is enforced on runs instead of dollars:
- Wall-clock timeout per stage —
implement30 min, all other stages 10 min. Passed towithAgentTimeout()at each call site. - Run-count cap —
writeRunLog()increments a per-stage counter on everyrun_startedevent and throws onceCURSOR_MAX_RUNS_PER_STAGE(default 20) is exceeded. Catches tool-call iteration loops. - Agent-creation cap —
withAgentTimeout()throws once a stage has calledAgent.create()more thanCURSOR_MAX_AGENTS_PER_STAGE(default 3) times. Catches retry loops.
Counters are per-process. Each npx run-stage invocation starts fresh.
When $GITHUB_STEP_SUMMARY is set, each stage appends a Stage | Duration | Cost | Runs | Status row (Cost shows n/a) so reviewers see wall-clock + run count without opening the Actions log.
Hooks
Seeded into the target repo by cursor-bootstrap. Registered in hooks/hooks.json.
| Hook | Event | Purpose |
|---|---|---|
| block-destructive.js | beforeShellExecution | Deny destructive shell commands before they execute |
| guard-mcp.js | beforeMCPExecution | Per-stage GitHub MCP allowlist + audit log of every MCP call |
| approve-issue.js | beforeMCPExecution (create_issue) | Step-up approval before an issue is created |
| approve-pr.js | beforeMCPExecution (create_pull_request) | Step-up approval before a PR is opened |
| log-edit.js | afterFileEdit | Append every file edit to .cursor-pipeline-edits.log |
block-destructive.js sources patterns from automations/shared/safety-denylist.ts — the same list Claude SDK's hook consumes. Blocks include rm -rf ~/.ssh, cat .env, curl … | sh, gh secret set, git filter-branch, --no-verify, force-push, and fork bombs. The local copy at hooks/safety-denylist.mjs is regenerated via npm run prebuild; do not edit it.
guard-mcp.js restricts which GitHub MCP tools each stage may call. Write stages are pinned to a single tool (create_issue → create_issue, create_pr → create_pull_request); read-only stages (review) may call get_/list_/search_ tools but no mutation; create_pr_draft may call none. The stage arrives via CURSOR_PIPELINE_STAGE (set by run-stage.ts / review-pr.ts). Every call (allow and deny) is appended to .observability/tool-calls.jsonl with the stage, traceId, tool, and decision. Override the path with TOOL_AUDIT_LOG_PATH.
Cloud runtime caveat: .cursor/hooks/*.js only fire under local runtime. buildAgentConfig throws when cloud runtime is selected without CURSOR_ALLOW_CLOUD_UNGATED=1 — set it per invocation to acknowledge that destructive-shell blocking, PR/issue approvals, and edit logging do not apply.
Note: guard-mcp.js governs the agent's own GitHub MCP tool calls during coding runs. The separate gh CLI calls in stages/create-pr.ts, fix-pr-comments.ts, and review-pr.ts are deterministic script code outside the agent loop (fixed API calls, not agent-decided) — execFileSync('gh', …) is the right tool there, not MCP. automations/shared/mcp-github-server.ts exists as an opt-in structured-tool-call alternative for stages that do need it (see the Claude SDK README, where no MCP-based GitHub layer exists yet).
Input sanitization
sanitize.ts (autogenerated from automations/shared/sanitize.ts via npm run prebuild) guards user-controlled text before it enters any LLM prompt.
sanitizeUntrustedInput({ text, source }) — for free-form user content:
- Strips
<tool>,<function_calls>,<invoke>,<result>,<param>XML markup (tag content preserved). - Checks 16 denylist patterns — instruction overrides, role-change attacks, jailbreak keywords, system-tag injection. Returns
[TRUNCATED: <source>]and logs[SECURITY]on match. - Truncates at 4 000 chars and wraps clean text in
<UNTRUSTED_USER_INPUT source="…">boundary tags.
sanitizePipelineHint(value, source) — for short workflow metadata (type hints, title hints, platform names):
- Strips newlines, caps at 200 chars, applies the same denylist. Returns empty string on a match (the hint is silently dropped from the prompt).
Currently wired into: triage (WORK_ITEM, STORY_TYPE_HINT, PIPELINE_TITLE, each PIPELINE_PLATFORM), create_spec (ISSUE_TITLE, ISSUE_BODY), implement (ISSUE_URL, TRIAGE_ISSUE_TITLE, TRIAGE_SUMMARY, TRIAGE_COMPONENT), review-pr (PR_URL, WORK_TYPE), fix-pr-comments (PR_TITLE, PR_DESCRIPTION).
Rule of Two. The pipeline simultaneously processes untrusted free-text (work_item, issue bodies, PR comments) and holds write-capable secrets and PR-creation ability — the combination the industry calls out as the risky one for agentic workflows. sanitizeUntrustedInput/sanitizePipelineHint above are the mitigation in place today; treat any new stage that reads untrusted text as needing the same sanitization before it reaches a prompt.
Scripts
| Script | Purpose |
|---|---|
| run-stage.ts | Stage driver: triage, create_issue_draft, create_issue, create_spec, implement, create_pr_draft, create_pr |
| build-implement-matrix.ts | Emits Actions matrix JSON from project marker file detection |
| detect-platform.ts | Resolves the platform slug from CURSOR_PLATFORM or marker file scan; sets platform_slug output |
| bootstrap-from-ticket.ts | Initialises pipeline state from a GitHub issue, Linear ticket, or Shortcut story URL |
| review-pr.ts | PR reviewer; structured verdict + deletions recommended; posts inline comments via gh |
| fix-pr-comments.ts | Triages unresolved review threads, applies fixes, posts inline replies, resolves threads. Phase 3 (replies + resolution) is skipped when Phase 2 produces no staged changes — no commit SHA to reference, and skipping prevents spin on the next invocation. Phase 4 (PR metadata normalization) always runs. |
| lib/knowledge.ts | Merges knowledge/shared/ with project .cursor/knowledge/ overrides into the ## Project Knowledge prompt block |
Observability
Cost log: .observability/costs.jsonl (gitignored). Records stage name, duration, success/failure, and traceId per run. Token and cost fields are always 0 — Cursor SDK does not expose model pricing. Configured in config/cost-tracker.ts.
traceId is a UUID generated once at pipeline start and written into every cost entry for that invocation, so all stage costs for one pipeline run can be grouped and queried together. Set via setTraceId() in lib/agent-utils.ts; call it once after state is loaded or created (already wired in run-stage.ts and bootstrap-from-ticket.ts).
Summary: npx cursor-sdk-cost-report
Every agent-running CI job uploads .cursor-pipeline-runs.jsonl and .observability/costs.jsonl as an observability-<runId>-<stage> artifact (30-day retention, if: always()). Download from the Actions run page to debug cost spikes, timeouts, or loop events post-mortem.
Publishing (maintainers)
Driven by sdk-publish.yml, triggered on a published GitHub Release with a cursor-sdk/vX.Y.Z tag.
One-time setup
secrets.NPM_TOKEN— npm automation token with publish access to the@nimblehqscope.- The
@nimblehqscope must exist on npm and the publishing user must havedeveloper/owneraccess.
Procedure
- Bump
versioninpackage.jsonand add a matching entry toCHANGELOG.md. Merge tomain. - Run Create GitHub Release → select
cursor-sdk. Version is read frompackage.json— no input needed.
The workflow reads the version from automations/cursor-sdk/package.json, creates the cursor-sdk/vX.Y.Z tag, publishes the release, and triggers sdk-publish.yml to build and push to npm and GitHub Packages. sdk-publish.yml re-checks that the tag matches package.json before publishing.
Diagnostics
| File | Contents |
|---|---|
| .cursor-pipeline-state.json | Pipeline state passed between stages (also uploaded as an artifact) |
| .cursor-pipeline-runs.jsonl | Agent and stage run lifecycle events (also uploaded as observability artifact) |
| target-repo/.cursor-pipeline-edits.log | Append-only edit audit written by hooks |
| .observability/tool-calls.jsonl | Per-MCP-call audit log: stage, traceId, tool, allow/deny decision (also uploaded) |
| .cursor-pipeline-review-raw-<ts>.txt | Raw agent output dumped by review-pr.ts when JSON parsing fails |
