npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@trybullet/cli

v1.4.12

Published

The fastest coding agent — terminal edition. Same router, tools and agent loop as the Bullet desktop app, in your shell.

Downloads

2,965

Readme

Bullet ⁍

The fastest coding agent. A minimalist macOS desktop app: open a folder and Bullet plans, edits files, runs commands, verifies its work, and shows you everything live. Connect a cloud provider or explicitly download the optional on-device model in Settings.

How it works

  • Agent loop (custom, ~200 lines — no framework): plan → tool calls (read / edit / write / delete / run) → observations → verify → done. Hardened with failure-retry directives, stuck detection, duplicate suppression, a 45s wall-clock cap, and a definition-of-done verification gate.
  • Router: a nano-model semantic classifier (hard 4s cap — routing is never the wait) rates each prompt on five difficulty levels. The optional local route runs Ministral 3 8B via Metal after an explicit Settings download. With a key → each level buys a deeper model and more thinking effort: quick asks run on fast models (Haiku / Codex-low), harder work escalates (Sonnet 5 / Codex at rising effort), and extreme tasks go to Fable 5 / Sol at full effort (say "use fable", "use sol", or "think hard" to force the level). The debug badge shows the exact model, tier, and effort used.
  • Plan gate: asking to plan ("help me plan X", "let's figure out the approach first") is its own classifier answer — those turns draft a plan and stop instead of working. Difficulty alone never triggers it: a big task you told Bullet to do just gets done.
  • Providers: Bring a ChatGPT/Codex or Claude Code subscription, or use an Anthropic, OpenAI, or xAI API key. Switch methods and providers in Settings (⌘,). OAuth credentials are encrypted locally; keys stay on this Mac.
  • Tools: read / write / edit / delete, search_code, find_files, run_command, background processes, web search + fetch, and a real browser (below). File tools reach any absolute path; commands run from the workspace.
  • Remote SSH workspaces. Open user@host:/path/to/project — from the sidebar (Connect over SSH…) or bullet -C user@host:/path — and Bullet works on that machine. Transport is your own ssh binary, so ~/.ssh/config, IdentityFile, ssh-agent, ProxyJump and Host aliases (Daytona, Tailscale, etc.) all just work; one multiplexed ControlMaster connection carries every command. The remote folder is mirrored locally (dependency trees excluded), file tools edit the mirror — diffs, snapshots and rewind included — while every run_command/start_process executes on the host: edits rsync across before each command, and whatever the command changed (lockfiles, generated code, git commits) syncs back after it. When a remote dev server prints a localhost URL, its port is forwarded over the same connection so Preview and the agent's browser reach it. Hosts without rsync fall back to tar-over-ssh. Keys, ssh-agent, password and key-passphrase prompts, first-use host confirmation, and keyboard-interactive/MFA are supported; secrets go directly to OpenSSH and are never saved. A one-off key file (a cloud provider's -i key) can be set in the SSH input or with --ssh-key <file>; it is remembered per host. Use ssh://user@host:2222/path for non-standard ports.
  • Code search that lands. search_code is extended-regex, so (a|b), \d and + mean what the model thinks they mean — grep's default dialect treats them as literal characters, which silently answered "no matches" for code that was right there. A pattern that finds nothing is retried as literal text before that answer is allowed, matches inside minified files are windowed with a column instead of dumping a 400 KB line, *.{ts,tsx} is expanded, and a directory you explicitly point at is never excluded as noise.
  • Eyes. open_pagescreenshotread_pagepage_action / page_eval drive an offscreen Chromium (no Playwright, no extra dependency). The screenshot goes to the model and into the transcript, so a visual change is verified by looking at it. read_page always reports console errors and failed requests, every load clears the cache, and page_action includes relative mouse-look for pointer-locked pages. Older captures are pruned from the context so a long visual session doesn't pay for stale pixels.
  • Speed levers: warm KV sessions (no re-processing between turns), patch-based edits (fewer output tokens), prompt caching on Anthropic, parallel tool execution on cloud, flash attention on Metal, idle model unload (frees ~2 GB after 5 min). One dev server per command: re-running npm run dev returns the server already running with its URL rather than starting a seventh one on a seventh port, and stop_process ends it.

UI

  • Ask (⌘/ or /ask …) — a side question in a floating popup, answered next to the run instead of interrupting it. Nothing it does can break the work in flight: the popup has its own thread and its own chat id, the answer never enters the transcript, and the main process strips every mutating tool (write/edit/delete, commands, processes, browser, sub-agents) from the turn — so Ask can read the project and the web and nothing else. It sees the recent turns of the chat it was opened over, so "why did it do that?" has an antecedent. Drag it by the header, Esc (from inside) closes it, and closing keeps the answer streaming for when you reopen.
  • Sidebar grouped by folder — chats are bound to their folder; agents in different folders run concurrently (busy dot), and agents in the same folder are told about each other's recent edits.
  • Live activity: plan checklist ticks in real time, action rows (✓ Edited app.js +10 −3, click to expand output), a status pulse with an elapsed counter through every silent reasoning stretch (never looks frozen; dead connections error out via stream watchdogs), and a stop button (Esc).
  • Resizable panels: drag the sidebar and Activity/Preview edges; widths persist.
  • Activity panel (right): running/finished commands + per-file diffs.
  • Plan panel: a plan-first request drafts markdown (goal, approach, risks, steps) into the side panel and waits — reply with feedback to iterate, ask questions and the plan stays put, edit it directly, then Run plan (or "run it") executes. Steps tick as the agent works; mid-run edits steer it. The plan lives in the panel, never in the chat. Discard, or just move on to other work, exits plan mode.
  • LaTeX: $$…$$, \[…\], \(…\), and math-y $…$ render via KaTeX (bundled, offline); dollar amounts in prose stay literal.
  • Preview panel: in-app browser for whatever the agent is serving. It opens itself when the agent starts a dev server, and auto-reloads (cache cleared) after each edit burst — toggleable via the auto button.
  • Prompt history: ↑/↓ in the composer cycles your previous prompts, terminal-style; your in-progress draft is preserved.
  • Attachments: 📎, drag in, or paste code/text, images & screenshots (sent as vision to cloud models), and PDF / Word / Excel / PowerPoint (auto-extracted to text — works with the local model too). Rewind/fork: hover any of your messages — rewinding confirms first, then also reverts the file edits made after that point (command side-effects are not undone). Two themes: Midnight / Paper.

CLI

The terminal front end. Not a port — it loads electron/router.js and every module under it, so routing, tools, skills, sub-agents, the rewind ledger and the provider auth are byte-for-byte the same code the app runs.

npm link                          # puts `bullet` on your PATH
bullet account signup             # create a Bullet account (first run also offers this)
bullet account guest              # start without an email; upgrade later without losing history
bullet providers                  # connect or inspect providers and local servers
bullet auth openrouter            # connect a specific provider directly
bullet auth status                # show every connection
bullet                            # interactive session in this folder
bullet "fix the failing tests"    # run one task, then keep talking
bullet -p "what does app.js do"   # print and exit — pipeable
git diff | bullet -p "review this"

Inside a session: enter sends, ctrl+j newlines, esc stops a running turn, tab completes /commands and @files, walks history. /help lists the rest — /providers, /model, /thinking, /cost, /diff, /undo, /resume, /agents. /login signs in to your Bullet account; /providers connects and switches model providers.

Providers and thinking

bullet providers (or bullet auth) presents the full provider surface: Claude and ChatGPT/Codex subscriptions; Anthropic, OpenAI, xAI, Gemini (Google AI Studio), Groq, and OpenRouter API keys; Ollama and LM Studio local servers; and a custom OpenAI-compatible endpoint. Each connection keeps its own encrypted key, endpoint, and selected model, so adding one never replaces another.

Provider presets are used when you press Enter at the base-URL prompt:

bullet auth gemini       # Google AI Studio
bullet auth groq
bullet auth openrouter
bullet auth ollama       # http://localhost:11434/v1
bullet auth lmstudio     # http://localhost:1234/v1
bullet auth custom       # enter an HTTPS endpoint
bullet auth status

Endpoints without a scheme are saved as HTTPS. HTTP is accepted only for loopback local servers. Local servers do not require a key. Use a provider and thinking level per run or per session:

bullet --provider openrouter --model openai/gpt-5.2 --thinking high "review this"
bullet --provider ollama --model qwen3-coder "fix the tests"
# in a session — unconfigured providers connect before becoming active
/providers gemini
/model gemini-2.5-pro
/thinking high            # auto | low | medium | high | xhigh | max

--effort remains an alias for --thinking. The setting is persisted in the session and forwarded only where the selected model supports it: OpenAI reasoning models, Claude, supported xAI/Gemini/Groq models, and OpenRouter models that advertise reasoning metadata. Unsupported local and compatible models receive no reasoning parameter, so the setting does not cause API errors.

Skills come across from Claude Code and Codex without the desktop app:

bullet skills import              # review what's in ~/.claude and ~/.codex, tick what you want
bullet skills import ./team-kit   # scan another folder too
bullet skills import --all        # no review — for scripted setups
bullet skills export kit.json     # move your reviewed skills to another machine

SKILL.md files, commands/*.md (including nested /review:pr), agents/*.md and CLAUDE.md are all recognised, and nothing is stored until you select it. /skills does the same thing inside a session, plus enable, disable and remove. Imported commands then run as slash commands and complete on tab; built-ins always win, so an import can't shadow /undo. Skills that depend on external MCP servers still need those tools configured separately; importing a skill does not make its MCP dependencies available in Bullet.

MCP connections work in the CLI too. Import one already configured in Codex or Claude Code, add any remote HTTP or local stdio server, then test it before use:

bullet mcp list                         # connections and current status
bullet mcp import linear                # import a discovered Codex/Claude connection
bullet mcp add my-api https://example.com/mcp
bullet mcp add local-tool -- npx -y @example/mcp-server
bullet mcp search datadog                # includes registry IDs
bullet mcp install com.mcparmory/datadog # guided encrypted setup
bullet mcp configure datadog            # update an existing connection
bullet mcp enable linear
bullet mcp login linear                 # browser OAuth when status is needs-auth
bullet mcp test linear                  # connect and list the tool count

Use bullet mcp search <query> to search the MCP Registry, bullet mcp help for the complete command list, or /mcp for the same workflow inside an interactive session. OAuth tokens and any values supplied with --header or --env are encrypted in Bullet's credential store. mcp install and mcp configure prompt for setup fields in an interactive terminal; scripts can pass them explicitly, for example bullet mcp configure datadog --env API_KEY_AUTH=… --env APP_KEY_AUTH_API_KEY=….

Model-provider sign-in reuses whatever is already on the machine: an existing Claude Code or Codex login is imported directly, no browser round-trip. Credentials are encrypted with the OS keychain (macOS Keychain, secret-tool on Linux) and kept in ~/.bullet, separate from the app's own store.

Guest mode is a real, device-local account: its chats and usage are attributed on the dashboard, but it has no email recovery. Use bullet account upgrade or Settings → Account to add an email and password without losing that history.

In scripts, -p exits 0 when it answered, 1 when the turn errored, 130 when you interrupted it, and 2 on bad arguments — so bullet -p … || handle-it does what it looks like. stdout carries only the answer; routing notes and warnings go to stderr.

Two capabilities do not cross over. The agent's browser — open_page, screenshot, read_page — needs the offscreen Chromium only Electron can host, so those tools are removed from the kit and the system prompt is told, rather than left to fail mid-turn. And with no provider connected the CLI falls back to the same Ministral 3 8B model and download path the app uses; that is a one-time ~5 GB download, so it asks first, and -p refuses rather than pulling it in a script nobody is watching.

How it stays one codebase: cli/shim.js registers itself as the electron module before anything under electron/ is required, and supplies the four things the shared core actually asks Electron for — a user-data directory, safeStorage, shell.openExternal, and a BrowserWindow that explains why it isn't there. Everything in cli/ above that line is presentation.

Run

npm install
npm start        # dev
npm test         # UI liveness + feature regressions (offline, no key)
npm run cli      # the CLI, from source
npm run test:cli # CLI suites (plain node, offline, no key)
npm run dist     # local unsigned build (right-click → Open)

npm run eval:classifier additionally checks routing quality — plan intent and plan-approval verdicts — against a live nano model, so it needs a provider key.

Linux

Every host difference resolves in one place, electron/platform.js; nothing else in electron/ or src/ may hardcode a path or read process.platform, and scripts/check-platform.sh fails the build if it does. CI runs the suite on macOS and Ubuntu both.

npm run dist:linux            # AppImage + deb into release/
npm run release:publish linux # push to the linux update feed

Which package to hand a user

| | .deb | AppImage | |---|---|---| | Ubuntu / Debian | recommended | needs setup, see below | | Other distros | — | recommended | | Updates itself | no | yes |

The split is not a preference, it is what works. Chromium's sandbox needs either unprivileged user namespaces or a setuid helper. Ubuntu 24.04 blocks the former (kernel.apparmor_restrict_unprivileged_userns=1) and an AppImage is mounted ro,nosuid, so the helper can never be setuid there — the app aborts on launch with "The SUID sandbox helper binary was found, but is not configured correctly." A .deb can make the helper setuid, which build/deb-postinst.sh does; read that file before touching it, because electron-builder's generated postinst gets this wrong in a way that only breaks for the user, not the packager. Ubuntu 24.04 also ships FUSE3, so an AppImage needs libfuse2t64 before it will even mount.

Running with --no-sandbox would paper over all of it and is deliberately not done: the Preview panel renders pages the agent navigated to, which is exactly what that sandbox is there to contain.

Two Linux gaps to know about

  • No on-device model. node-llama-cpp kills the process while loading a model inside Electron on Linux — SIGILL on Electron 32, SIGSEGV on 34 and 38 — while the same model and binaries work under plain node on the same machine. A native crash cannot be caught, so platform.localEngine() gates it off and onboarding asks for a subscription or API key instead. Fixing it means running the engine in a child node process and streaming over IPC.
  • .deb cannot self-update. electron-updater only supports AppImage on Linux, so Settings tells .deb users where to download from rather than offering a button that cannot work. A real fix is an apt repository.

The AppImage ships CPU and Vulkan llama backends only; the CUDA ones are excluded in build.files because they are 526 MB and load solely where the CUDA runtime is already installed, which for most desktops means never. Vulkan already covers NVIDIA, AMD, and Intel. Re-adding them is one glob.

Windows

Windows ships as a 64-bit NSIS installer and updates itself through the Windows update feed.

npm run dist:win                # installer + latest.yml into release/
npm run release:publish windows # push to the Windows update feed

Agent commands run through Git Bash, which is installed with Git for Windows; Bullet reports a clear startup error instead of accidentally invoking WSL when Git Bash is unavailable.

Windows builds are deliberately unsigned: a publicly trusted Authenticode identity is not available for free, and a self-signed certificate would still show a security warning while providing no trustworthy proof of identity. The installer therefore shows Windows' honest Unknown publisher warning, while its filename, icon, file metadata, installed-app entry, HTTPS download URL, and SHA-256 file published beside every GitHub release identify it as Bullet. Users should download only from the official Bullet site or trybullet GitHub release. No Windows signing secrets are required by the release workflow.

The on-device model is currently disabled on Windows; subscriptions and API keys work normally.

Disposable test installers and cloud desktops

The Test installers workflow builds an unsigned Apple Silicon DMG, Windows NSIS installer, Linux AppImage, and Ubuntu/Debian .deb without publishing any of them:

npm run test:installers

Download the three seven-day artifacts from the workflow run in GitHub Actions. They are suitable for Google Drive and manual testing. macOS requires right-clicking the app and choosing Open because the test DMG is unsigned; Windows shows Unknown publisher by design. On Ubuntu, use the .deb; use the AppImage for other compatible x64 distributions.

Disposable Windows 11 Pro x64 and Ubuntu 24.04 x64 Azure desktops can be created and removed from a machine where Azure CLI is logged in:

npm run test:vm:up          # create Windows; print RDP credentials
npm run test:vm:down        # delete the Windows resource group
npm run test:linux-vm:up    # create Ubuntu Desktop; print RDP credentials
npm run test:linux-vm:down  # delete the Linux resource group

Linux is based on Canonical's official Ubuntu 24.04 server image, with the minimal GNOME desktop and xrdp installed automatically. Use Microsoft Windows App for either desktop. On Ubuntu, download and install the .deb artifact; that is the recommended package and exercises Chromium sandbox setup too.

Creation permits RDP only from the caller's current public IP and enables a UTC auto-shutdown (0100 by default). Override settings when needed, for example:

BULLET_TEST_VM_SHUTDOWN=2330 BULLET_TEST_VM_LOCATION=eastus npm run test:vm:up
BULLET_LINUX_TEST_VM_SHUTDOWN=2330 npm run test:linux-vm:up

Auto-shutdown deallocates compute but does not delete storage; always run the matching down command after testing. The Windows and Linux desktops use separate resource groups and can exist at the same time.

Releasing

Installed copies update themselves: the app checks https://api.codewithbullet.com/bullet/v1/updates/mac/latest-mac.yml on launch and every six hours, downloads in the background, and installs on quit (Settings › Updates has a manual check and a Restart to update button). Nobody has to delete and reinstall.

This only works for signed builds. macOS applies updates through Squirrel.Mac, which requires the downloaded bundle to satisfy the running app's code-signing requirement — an ad-hoc signed build can never update itself, and publishing one strands every user who installs it on that version. scripts/publish-update.sh refuses to publish an unsigned build for that reason.

One-time setup on the build Mac:

  1. In Xcode › Settings › Accounts, sign in with the Apple Developer account, then Manage Certificates › + › Developer ID Application. Confirm with security find-identity -v -p codesigning — you need a line reading Developer ID Application: … (TEAMID), not just Apple Development.

  2. Create an app-specific password at https://appleid.apple.com › Sign-In and Security › App-Specific Passwords — notarization rejects the account password — and put the credentials in .env.notarize at the repo root:

    APPLE_ID="[email protected]"
    APPLE_APP_SPECIFIC_PASSWORD="xxxx-xxxx-xxxx-xxxx"
    APPLE_TEAM_ID="YOURTEAMID"   # Membership details on developer.apple.com

    npm run release loads that file itself. It is gitignored, it must stay that way, and the Team ID in it has to match the team that issued the signing certificate — the script checks and stops if they disagree, because Apple rejects the notarization otherwise.

Then each release:

npm version patch          # or edit "version" — the feed compares against it
npm run release            # signs + notarizes, writes release/latest-mac.yml
npm run release:publish     # uploads to the update feed on the API box

npm run release takes a few extra minutes the first time — notarization is a round trip to Apple. Running clients pick the new version up within six hours, or immediately from Settings › Updates.

Signing on someone else's Mac

Bullet is signed on one specific Mac, and that never changes. Squirrel.Mac only applies an update whose signature satisfies the running app's requirement, so every release must be signed by the same Developer ID / Team ID as the one before it. Signing a release under a different team strands every installed copy — those users would each need a manual reinstall to get updates back. Keep a .p12 export of that certificate (and its password) somewhere safe off that machine: losing access to it means losing the ability to ship updates to everyone already running Bullet.

Whoever signs builds from this repo rather than re-signing someone else's .app — electron-builder rebuilds the native modules and signs the unpacked node-llama-cpp binaries as part of packaging, and any hand-editing of a signed bundle invalidates both the signature and the sha512 in latest-mac.yml.

On their Mac: Xcode command line tools, Node, the Developer ID Application certificate in the login keychain, then

git clone … && cd bullet && npm install
# write .env.notarize with the Apple ID, app-specific password, and the Team ID
# that issued the certificate on this Mac (see step 2 above)
npm run release

(CSC_NAME="Developer ID Application: Name (TEAMID)" if the keychain holds more than one Developer ID certificate.) Leave build/entitlements.mac.plist alone — JIT, unsigned executable memory, and library validation are all off-limits to strip, because the on-device model needs them under the hardened runtime.

Checks before handing the build over:

codesign -dv --verbose=2 release/mac-arm64/Bullet.app  # Authority=Developer ID Application: … (TEAMID)
spctl -a -vv release/mac-arm64/Bullet.app              # accepted, source=Notarized Developer ID
xcrun stapler validate release/Bullet-*-arm64.dmg

Then hand over release/ intact — the .zip (electron-builder's, never a Finder-made one, which mangles symlinks and the signature), its .blockmap, the .dmg, and the latest-mac.yml from that same build, whose hashes must match the files beside it. npm run release:publish does the upload and needs the VM SSH key, so it is easiest for whoever holds that key to run it on the received release/ directory.

SWE-bench Verified

scripts/swe/ runs Bullet's real agent over SWE-bench Verified and grades every patch with the official Docker scorer — see scripts/swe/README.md.

All 500 instances, officially scored: 479/500 = 95.8%. Native x86, one attempt per instance, network on, gpt-5.6-sol, swebench.harness.run_evaluation. Zero empty patches, zero worker errors, every instance graded.

| | resolved | round trips¹ | latency | cost/inst | |---|---|---|---|---| | Bullet — all 500 | 479/500 = 95.8% | 15.0 | 119s | $0.73 | | mini-swe-agent + sol@max — all 500 | 421/500 = 84.2% | 12.6 | 91s | $0.29 |

Resolve rate by repository, so it is clear no single project is carrying it: django 96.5% (223/231), sympy 97.3%, sphinx 95.5%, matplotlib 94.1%, scikit-learn 100%, xarray 95.5%, astropy 81.8%, pytest 100%, pylint 90%, requests 87.5%.

The 17-point asterisk. Both rows above have network access, which is the reference's default and — on the evidence of its cost and latency — the published leaderboards' too. That matters enormously: these instances are public GitHub PRs, so an agent with a route out can fetch the merged upstream fix, and the reference does exactly that on 23% of instances (curl raw.githubusercontent.com/<the very file>, api.github.com/search/issues?q=<the bug>). Measured directly on the random 120, same configuration, only the network changed:

| | resolved | |---|---| | network on | 113/120 = 94.2% | | --network none, no web tools | 92/120 = 76.7% |

So 95.8% is the leaderboard-comparable number and ~77% is the from-the-repository -alone number. Offline is the default here; SWE_NETWORK=1 is the switch.

¹Round trips — model calls per instance — is the only speed figure worth quoting. Wall clock is provider-load dependent: the same configuration on the same instances measured 131s and 248s nine hours apart with the machine idle. Turns don't drift. profile.js reports both.

What the turn-efficiency work bought, paired on the same instances: 17.2 → 14.4 round trips and $0.93 → $0.68 an instance, for an accuracy change that is not statistically separable (0 gained, 2 lost, McNemar p = 0.50 — a 1-2 point cost is possible and unproven). Against the reference on the shared 120, Bullet resolved every instance it resolved and 17 more, nothing the other way (p < 0.001).

A gold control on this machine resolves 117/120 reference patches on the random 120; the three exceptions (a live HTTP call, an env-sensitive header test, an extension-version test) cannot pass whatever the patch does. So roughly 2.5% of any raw rate here is the machine, not the agent — ceiling.js reports the corrected figure. No gold control was run over all 500, so the 95.8% above is raw.

What moved it. Routing: the deep model resolves 62.1% at 138s against the fast model's 36.8% at 152s, so the best model leads at every level instead of being an escalation. Reliability: an earlier run submitted 13 blank patches from dead workers, and there are none now. Two defects found by measurement rather than by reading:

  • search_code ran in grep's BRE dialect, where (, | and + are literal characters — so function (alpha|beta) reported "no matches" for code plainly present. Searches are 40% of all tool calls, so this cost both accuracy and a great many wasted round-trips.
  • The shared git mirror had a creation race. git clone makes its directory immediately and fills it over minutes, so a worker that only tested existence cloned from a half-written mirror and lost the instance one second in. Six of 120, all in the repositories eight workers wanted at once; all six resolved on re-run.

A third defect, found by a one-instance smoke test after the runs above: an undeclared timer variable threw at the first phase write, silently aborting the tail of every attempt — losing the reproduction it was meant to collect, which is what the candidate selector runs on. The patch survived because the error path recovers it from disk, which is exactly why it went unnoticed.

Two other levers were paired, graded and landed within noise: the methodology brief versus bare environment facts, and structured tools versus a bare shell.

Latency is round-trip count and nothing else. Instrumented per tool over the 500: an instance is 120s, of which checkout 0.3s, container start 0.7s, collect 0.5s, tool execution 14s (12%, four fifths of it real test suites) — and 104s of model thinking, i.e. 15 round trips at 7s each. There is nothing to win in the scaffolding; the only lever is needing fewer turns. Bullet's 15 against the reference's 12.6 is the entire speed difference between them. Beware comparing tool calls instead: Bullet issues ~38 of those, but batches 2.5 per turn, and confusing the two overstated the gap threefold in an earlier draft of this file. profile.js reports both.

Two things carry the result, and a third had to be taken back out:

  • A testbed the agent can run. The attempt works in the instance's prepared container, so it can actually reproduce the bug and run the project's tests instead of reasoning about code it cannot execute. File tools keep editing a local checkout (that is where the diff comes from) and edits are pushed across before every command.
  • The best model, at full depth, first. Measured over 86 graded attempts, the deep model resolved 62.1% at 138s against the fast model's 36.8% at 152s — better and quicker, because it needs far fewer steps. Escalation is right for chat and wrong here, so there is no ladder: one attempt, strongest model, highest effort.
  • The ensemble is off by default, and the reason is not yet settled. Several attempts under different lenses, each leaving a reproduction, every reproduction run against every candidate in fresh containers, a reviewer to break ties — that configuration scored 104/120 against a single deep attempt's 113/120, losing nine instances and winning none (p = 0.004). But that comparison is confounded: an undeclared variable in the phase stopwatch was throwing partway through every attempt, and the statement it aborted was the one that collected the attempt's reproduction. So the selector — whose whole job is weighing reproductions — was running on almost none, and fell back to preferring the first, shortest candidate on 97 of 114 instances. The bug is fixed; the experiment has not been repeated. One attempt is the default because it is the best measured configuration, not because the ensemble is known to be worse. An oracle over the pair scored 105/114, so the headroom is real and SWE_K still turns it back on.

The benchmark's answer key — test patch, reference patch, FAIL_TO_PASS/PASS_TO_PASS — never reaches inference; export-instances.py asserts those fields out of the data the agent sees. SWE_NETWORK=1 is the switch behind the two regimes in the table; without it the web tools are removed and every container runs --network none.

First prompt downloads the on-device model once (~5 GB), then it's local. Apple Silicon only.

Layout

electron/  main.js (window/IPC) · preload.js · router.js (routing+context)
           engine.js (local agent loop) · cloud.js (cloud loops) · tools.js
src/       index.html · styles.css · app.js
cli/       bin.js · shim.js (fake `electron` for plain node) · main.js (repl)
           session.js (host wiring) · render.js (events → transcript) · ui/
scripts/   icon/logo generators · ui-test.js / feature-test.js (offline suites)
           cli-test.js (CLI suites) · classifier-eval.js
           eval.js / swebench.js / polyglot.js (live evals)
BENCHMARK.md  one-shot website benchmark vs Claude Code / Codex