@babylonjs-toolkit/agent
v1.1.51
Published
Installs the Babylon Toolkit Desktop Agent for Claude Code, Codex, GitHub Copilot, Gemini CLI and Antigravity.
Maintainers
Readme
Babylon Toolkit Desktop Agent (1.1.51)
The desktop agent owns the entire pipeline end to end — frontend and UI design, gameplay code, shaders, generated art and audio, 3D models in headless Blender, whole game levels and prefabs in a terminal-driven Unity Editor, the interactive glTF export, the web build, the dev server, and visual QA by screenshotting both Unity and the running browser.
Universal Agent Skills for the Babylon Toolkit web game development framework.
Each SKILL.md follows the open standard, so the same file works unchanged in Claude Code, Codex CLI, and GitHub Copilot.
Install
npm install -g @babylonjs-toolkit/agentor, without installing anything globally:
npx @babylonjs-toolkit/agent installThat installs every skill into every skills directory and the Agent Persona into every global instruction file, on macOS, Linux and Windows. Codex targets also enable outbound network access in their workspace-write configuration. Restart your agent session to pick up the skills, instructions and configuration.
Commands
| Command | What it does |
|---------|--------------|
| bt-agent install | Install skills + persona into every target (the default) |
| bt-agent update | Fetch the newest release from npm, reinstall it, and prune dropped skills |
| bt-agent uninstall | Remove what it installed — and nothing else |
| bt-agent doctor | Verify the install; prints INSTALL OK or lists what is missing |
| bt-agent targets | Show every target and the paths it writes to |
| bt-agent bridge | Connect this computer to the App Builder so it can drive Unity and Blender |
| bt-agent diskinfo | Show where your disk space is going — read-only, macOS, Windows and Linux |
| bt-agent kill | Free a port or range of ports, stop processes by id, or list listening ports — macOS, Windows and Linux |
| Option | Effect |
|--------|--------|
| --project | Install into the current directory instead of your home directory |
| --targets claude,agents | Restrict to specific targets |
| --legacy-codex | Also write ~/.codex/skills for pre-.agents Codex builds |
| --no-persona | Install skills only; leave instruction files alone |
| --persona-only | Install/refresh the Agent Persona only; do not copy skills |
| --no-self-update | update only: reinstall the bundled files; do not fetch npm |
| --dry-run | Print what would happen and change nothing |
| --json | Machine-readable output |
Disk usage report
bt-agent diskinfo # print the report and save Desktop/disk-report-YYYY-MM-DD-HHMM.txt
bt-agent diskinfo --no-save # print only
bt-agent diskinfo --out report.txt
bt-agent diskinfo --json # machine-readableOne read-only pass over your home folder and the main system folders (/Applications, /Library,
/usr/local … on macOS; Program Files, ProgramData on Windows). It reports the volumes, the
largest folders, known space hogs (Unity editors and caches, Xcode DerivedData and simulators,
npm / Yarn / pnpm / pip / NuGet caches, Docker and WSL disks, VS Code's C++ cache, Steam, backups,
Trash / Recycle Bin), every node_modules, the Unity Library/ and Unreal Intermediate/ folders
that rebuild themselves, the largest files, and suggested cleanup commands. Nothing is deleted or run.
Freeing ports
bt-agent kill --port 4444 # stop whatever is listening on port 4444
bt-agent kill --port 4444 8888 # several ports (or 4444,8888)
bt-agent kill --port 8000-8010 # an inclusive range
bt-agent kill --pid 51234 # stop a process by id (or 51234,51240)
bt-agent kill --list # every listening port with its process id and name
bt-agent kill --list 3000-3999 # just these ports; stops nothing
bt-agent kill --list --pid 51234 # the ports that process listens on
bt-agent kill --port 4444 --force # skip the polite request and force-kill at once
bt-agent kill --pid 51234 --force # the same for a process id (-9 works too)Exactly one of --port, --pid or --list says what the values are, so a bare
number is never guessed to be a port or a process id.
Use it when a dev server is still holding its port and you need to restart it. With
--port, only processes listening on a port are matched, so a browser tab connected to your dev
server is left alone. On macOS and Linux each process gets SIGTERM, then SIGKILL if
it is still running after 3 seconds. On Windows the process and its children are
force-killed with taskkill /T /F. Afterwards the ports are checked again, and any
process a watcher (nodemon, pm2 …) has already restarted is reported. Ports are found
with lsof on macOS, ss (or lsof) on Linux, and netstat on Windows. A process
owned by another user needs sudo. bt-agent will not stop itself, the system init
process, or the Windows System process.
Updating
bt-agent updateupdate asks npm for the newest published version. If there is one it upgrades the
global package for you — with whichever manager installed it (npm, pnpm, yarn
or bun) — then hands over to the freshly installed CLI to copy the new skills,
refresh the persona, and prune skills the new release no longer ships.
If the upgrade is not possible it says so and continues with what is on disk:
- running through
npx, from a git checkout, anpm linked working copy, or as a project dependency — there is no global package to upgrade, so it reinstalls the bundled files - the registry is unreachable — it reinstalls the version you already have
- the global install needs elevated permissions — it prints the exact command to run
- npm reports success but this CLI still runs the old version — you have more than one Node/npm on PATH, and it tells you which path was left behind
Use --no-self-update to skip the npm check entirely.
What it writes
| Target | Skills | Agent Persona |
|--------|--------|---------------|
| Claude Code | ~/.claude/skills/ | ~/.claude/CLAUDE.md |
| Codex / Copilot / Antigravity | ~/.agents/skills/ | ~/.agents/AGENTS.md |
| Codex chat clients | — | ~/.codex/AGENTS.md |
| Gemini CLI | — | ~/.gemini/GEMINI.md |
| Legacy Codex CLI (--legacy-codex) | ~/.codex/skills/ | — |
Gemini CLI is a separate entry on purpose: it reads GEMINI.md, not AGENTS.md, so
~/.agents/AGENTS.md alone never reaches it.
Codex installs (the default codex target, or codex-legacy) also merge the shipped
.codex/config.toml setting into ~/.codex/config.toml, or
<project>/.codex/config.toml with --project:
[sandbox_workspace_write]
network_access = trueThis allows outbound network access to all domains while using the workspace-write sandbox, including fetching the Agent Reference from GitHub. It does not change the sandbox mode, filesystem permissions or approval policy. Managed policies can override this setting; project configuration applies only to trusted projects.
Install, update and the global npm postinstall hook apply this setting whenever a Codex
target is selected, including --persona-only and --no-persona runs. An existing
network_access = false is changed to true; other settings and comments are preserved.
Changed configurations are backed up to config.toml.bak; repeated installs are idempotent
and --dry-run writes nothing. Standard TOML sections and root dotted keys are supported;
an inline sandbox_workspace_write = { ... } table is left untouched with an actionable
error, so convert it to the section form above before retrying.
bt-agent doctor checks the setting, and bt-agent targets lists its destination.
Uninstall retains this configuration and its backup as user preferences. Selecting only
non-Codex targets leaves Codex configuration untouched.
Your instruction files are never truncated. The persona goes in as a managed block:
<!-- BEGIN BABYLON TOOLKIT PERSONA v1 -->
...
<!-- END BABYLON TOOLKIT PERSONA -->Re-running rewrites only what is between those markers, so the persona can be revised in a
later release while everything else in the file is left exactly as you wrote it. An older
unmarked persona is converted to a managed block automatically; a persona you have reworded
yourself is detected and left alone. Every modified file is backed up to <file>.bak first.
Skill folders are tracked in ~/.babylon-toolkit/install-manifest.json, so update and
uninstall only ever touch folders this package installed — a skill you wrote yourself that
happens to be named bt-something is safe.
Unity Bridge (App Builder)
bt-agent bridge connects this computer to the Babylon Toolkit App Builder, so the builder's
chat can open, create, edit and export Unity projects and drive Blender installed here.
Open the Unity Bridge dialog in the App Builder (the cube icon in the chat box), type your App
Builder projects folder (the one you picked for your projects in the App Builder), copy the one command
it then shows, and run it once in a terminal. The command points the helper at the Unity folder inside
your App Builder projects folder, where Unity projects live beside the Apps folder of apps (the
Unity Bridge dialog fills it in):
npx @babylonjs-toolkit/agent bridge --install-service --pair K7QM-2XWD --projects "/Users/you/Projects/Unity"--projects is required with --install-service — the service runs at login, so the folder is never
guessed from where the command happened to run. A folder that does not exist is created. A later
reinstall without --projects keeps the folders stored by the first one.
The command carries a single-use install code (valid for 10 minutes) that pairs this computer with
your account — you never type a code anywhere. --install-service then copies the helper to
~/.babylon-toolkit/service/ and starts it now and every time you log in (a launchd agent on
macOS, a systemd user unit on Linux, a Startup-folder script on Windows). It logs to
~/.babylon-toolkit/bridge.log (rotated at 5 MB). Run the command again to change the settings —
it re-installs and restarts the service. bt-agent bridge --uninstall-service stops it and stops
starting it at login; the pairing is kept until bt-agent bridge logout.
The App Builder can list the Unity projects in the projects folder, open one, or create a new one
there — the project it opened or created last is the one every other Unity and Blender job works on.
Run in the terminal (without --install-service) and without --projects, the bridge uses the folder
it runs in — or, started inside a Unity project, that project's parent, with that project open. Only folder and project NAMES are sent to the App Builder, never a path.
The helper talks to the production App Builder by default. --server <url> (or BTK_BRIDGE_SERVER)
points it at another one — the dialog adds it to the command for you when it is needed. --server
may be repeated to serve several App Builders at once (for example a local development server and
production): each is paired once, with its own install code, and they share one job queue and one
current project, so two builders never drive Unity at the same time.
| Option | Effect |
|--------|--------|
| --pair <code> | The install code from the Unity Bridge dialog |
| --install-service | Pair, then start the bridge now and at every login |
| --uninstall-service | Stop the service and remove it (the pairing is kept) |
| --server <url> | Another App Builder — https://, or http://localhost for development (repeatable) |
| --projects <folder> | A Unity projects folder (repeatable; created if missing) — with the App Builder, the Unity folder inside your App Builder projects folder (the Unity Bridge dialog fills it in). Required for --install-service; default only when running in the terminal: this folder, or its parent inside a Unity project |
| --unity <path> | Also serve this Unity project (repeatable); the first one starts as the current project |
| --blender <path> | The Blender executable to use |
| --no-scripts | Never run C# or Python scripts on this computer, whatever Allow scripts says |
Without --install-service, bt-agent bridge [--pair <code>] runs the bridge in the terminal until
Ctrl-C. With no stored pairing and no --pair, it says so and exits: copy the install command from
the Unity Bridge dialog. bt-agent bridge status shows the App Builders this computer is paired with
and whether the service is installed; bt-agent bridge logout [--server <url>] unpairs it (all App
Builders by default) and deletes the credentials. bt-agent doctor reports the service state too.
Credentials and service settings live in ~/.babylon-toolkit/bridge.json, readable by you only.
Everything the App Builder asks for falls into one of three tiers, and this computer enforces them itself:
| Tier | What it covers | When it runs |
|------|----------------|--------------|
| allowed | Listing, opening and creating projects in the projects folder; reading the project and ordinary edits | Straight away |
| scripts | Running C# or Python scripts | Only when Allow scripts is on for this computer in the App Builder's Unity Bridge dialog (on by default), and never with --no-scripts |
| consent | Deleting, moving, renaming, building, changing project settings | Only after you approve it in the chat |
The service is opt-in: bt-agent install, bt-agent update and the npm postinstall never start,
pair or schedule the bridge.
Skills
| Skill | Command | What it does |
|-------|---------|--------------|
| bt-spec | /bt-spec | Turn a short idea into a feature spec file on a new git branch. Add --grill-me to be interviewed first, --parity to allow numeric parity bars in the acceptance criteria. Runs in TIME MATTERS mode by default; --no-time-limit turns it off. |
| bt-plan | /bt-plan | Produce a detailed, task-checklist technical plan from a spec. Add --heavy for a decision-complete plan that keeps a long, fresh-context-per-task run cohesive, and --parity for numeric parity gates instead of the default functional proof. Runs in TIME MATTERS mode by default; --no-time-limit turns it off. |
| bt-execute | /bt-execute | Implement one task, a range (T3-T7, T12-, NEXT:3) or all remaining tasks from a plan/spec; with no task id it runs the next unchecked task. Add --auto-pilot for an unattended overnight run that never stops for human input, and --strict for an adversarial verifier on every task. Runs in TIME MATTERS mode by default; --no-time-limit turns it off. |
| bt-recon | /bt-recon | Deep-dive existing code — a subsystem named in a brief, or the whole codebase — and write an evidence-cited technical grounding spec (_specs/<slug>_recon.md) plus, when it has UI, a plain-language user guide (_specs/<slug>_guide.md). Built for taking over undocumented code; bt-spec and bt-plan read the recon as grounding. Includes an observe-only live pass when the brief gives a URL; --refresh updates a recon after code changes. |
| bt-convert | /bt-convert | Convert source code to Babylon Toolkit TypeScript. |
| bt-copycat | /bt-copycat | Re-create the specified website adapted to specified genre. |
| bt-landing | /bt-landing | Re-design the landing page, splash screen, preloader and custom overlays. |
| bt-gauntlet | /bt-gauntlet | Agent based gauntlet loop engineering — web games, Unity levels, or Blender models. (usage guide) |
| bt-prototype | /bt-prototype | Create any number of award winning frontend prototypes. |
| bt-design | /bt-design | Implement high quality frontend and in-game designs. |
| bt-hero | /bt-hero | Create smooth cinematic 3D scrolling hero sections. |
| bt-atlas | /bt-atlas | Generate texture atlas skin variations. |
Every tool derives the slash-command from the folder name (bt-spec/ → /bt-spec) and reads
the frontmatter name + description to decide when the skill applies. The allowed-tools
line is honored by Claude Code (auto-approves those tools) and safely ignored by Codex and
Copilot.
Heavy plans (/bt-plan --heavy)
A plan can hold 50+ tasks and run for days, with each task executed by bt-execute in a fresh
context. The only thing that survives between tasks is the plan file. --heavy makes bt-plan
write it as the feature's shared memory: a numbered ## Decisions log (every choice, its
rationale, what was rejected, which tasks it binds), a ## Design Reference (file map, interfaces
written as code, data shapes, algorithms with their constants, edge-case policy, verbatim
conventions to mirror, exact API usage, test strategy), and self-contained tasks that cite those
sections by name instead of relying on memory of earlier tasks. Before the plan is final, a fresh
read-only subagent audits it as a cold executor and every gap it finds is fixed. The result is
that task 40 uses the same names, shapes and patterns as task 3 — whichever session or day runs
it. Without the flag, bt-plan produces its standard plan; bt-execute runs either kind the
same way.
Proportional verification (/bt-execute default) and --strict
Execution time is dominated by verification, not by the code edits, so the loop is sized to the
feature. bt-spec records a size (small / medium / large); bt-plan uses it to set the task
budget (a small feature is 1–3 tasks in one phase), groups tasks into phases, marks each task's
Verify level (standard or live), and never pins absolute test counts. By default
bt-execute has the implementer write and run each task's named tests, then runs one
independent verifier per phase with a scoped charter (Acceptance, the Verify commands, the
tests re-run and judged for coverage, a diff review for real defects), always does live
Unity/browser QA on any task whose correctness shows only at runtime (rendering, visuals,
interaction, Unity↔Babylon parity) plus one end-to-end check at the end, and re-checks only
what failed.
Docs are read once by the orchestrator and handed to subagents as a brief instead of every
subagent re-fetching the Agent Reference.
/bt-execute --strict @_specs/<feature>_plan.md ALL
/bt-execute --auto-pilot --strict @_specs/<feature>_plan.md ALLProof is sized the same way. By default a feature is proven functionally: tests, cheap
reference-value fixtures where a reference exists, and for anything on screen a look at the
running result in the browser — against the spec and DESIGN.md, or beside a reference image
when there is one — judged by the independent verifier — with no ledgers, md5 records, archived
capture trees or standing gate suites unless the brief asks. Numeric parity gates (pixel
thresholds, repeats, every engine on every criterion) are opt-in: /bt-spec --parity … or
/bt-plan --parity …. Every plan ends with an ## Estimated execution time section that splits
build from prove hours, so you can see before running it where the time will go.
--strict restores full rigor: an adversarial verifier on every task (mutation checks,
oracles, re-deriving from the plan as it sees fit), live QA on every task with a rendered or
exported surface, a full fresh re-verify on each fix attempt, and a 5-attempt fix loop. It
combines with --auto-pilot. There is no separate testing subagent in either mode — the
implementer writes the tests and the independent verifier judges them.
Unattended runs (/bt-execute --auto-pilot)
/bt-execute --auto-pilot @_specs/<feature>_plan.md ALLAuto-pilot is for running a 50-task plan overnight. There is no one to ask, so bt-execute
owns the whole pipeline — code, UI, shaders, generated art and audio, Blender, Unity, export,
build, dev server and visual QA — and makes every call itself: it first looks the answer up in the
brief, plan, spec and SPEC.md, then the Babylon Toolkit Agent Reference (Unity → interactive glTF
→ BabylonJS script components), then the codebase, then senior-developer default — and logs each
one. It never asks a question and never ends its
turn until the requested range, or every task, is done; it stops early only for a missing plan, a
bad range, a stop point named in the brief, or you telling it to stop. Nothing is relaxed on quality: every task still has passing tests
and goes through an independent verifier, and a box only flips on a genuine PASS.
What changes is that a failure never halts the run: a failing task gets a bounded fix loop (3
attempts, 5 under --strict), is then marked ⏭️ DEFERRED (auto-pilot): <reason> in the plan, and the run moves
on; deferred tasks get one more pass at the end. Each verified task is committed as a checkpoint
on the current branch (an autopilot/<plan> branch is created if you are on main; nothing is
ever pushed), and a <plan>_autopilot.md run log next to the plan records every task outcome,
decision and deferral for you to read in the morning. Re-running the same command resumes and
re-attempts anything deferred.
One thing the skill cannot do is answer the host's tool-permission prompts — start the session in a non-prompting permission mode (Claude Code auto / bypass-permissions) and leave Unity, Blender and the dev server reachable if the plan needs them.
Universal Installations
Where each tool looks for GLOBAL skills (Default Installation — ~ = your home directory):
| Tool | macOS / Linux | Windows |
|------|---------------|---------|
| Claude Code | ~/.claude/skills/ | %USERPROFILE%\.claude\skills\ |
| OpenAI Codex | ~/.agents/skills/ | %USERPROFILE%\.agents\skills\ |
| GitHub Copilot | ~/.agents/skills/ | %USERPROFILE%\.agents\skills\ |
| Gemini CLI / Antigravity | ~/.agents/skills/ | %USERPROFILE%\.agents\skills\ |
On macOS
~expands to/Users/<user>/, on Linux to/home/<user>/, and on Windows%USERPROFILE%expands toC:\Users\<user>\.
Where each tool looks for PROJECT-LOCAL skills (only when the user explicitly requests a project install):
| Tool | Project skills directory |
|------|--------------------------|
| Claude Code | .claude/skills/ |
| OpenAI Codex | .agents/skills/ |
| GitHub Copilot | .agents/skills/ |
| Gemini CLI / Antigravity | .agents/skills/ |
Note: Install skills into BOTH the
.claude/skills/and.agents/skills/locations for maximum coverage —.claude/skills/for Claude Code,.agents/skills/for everything else (Codex, Copilot, Gemini CLI, Antigravity). By default use the GLOBAL (per-user home) locations; use the project-local locations only when the user explicitly asks for a project install. Copy whole folders into each — thebt-atlas/scripts/assets must be present in every location.
Babylon Toolkit Loop Engineering (Gauntlet Loop)
What it is in one line: you give it a brief and a quality bar; it interviews you, locks a target image and the cameras every screenshot is taken from, writes a loop card, then runs rounds of build → capture real evidence → fresh harsh critic scores a rubric and returns the full gap list → record until the success condition passes or a boundary fires — building with BabylonJS + the Babylon Toolkit, never Three.js. All state lives in _gauntlet/<name>/, so you can stop any time and resume days or weeks later.
Two deliverable kinds, one loop
The round protocol, critic contract, stall ladder and resume machinery are identical in every case. What changes is what the builder edits and how evidence is produced — settled in the interview and recorded in the loop card.
| Kind | The builder edits | Evidence is |
| --- | --- | --- |
| web-game | BabylonJS / Toolkit TypeScript, scene code, shaders, UI | the running game in a browser |
| unity | a Unity scene via the Unity CLI — terrain, light rig, reflection probes, bake settings, fog, tonemapping — with headless Blender as the tool for low-level model work (in place, so GUIDs survive) | a Unity camera snapshot rendered straight to PNG each round — the loop stays in Unity. The export + browser check is a deliberate checkpoint, and what you export (whole level, or one asset as a container) is your call at the time |
Each kind's prerequisite gate, part taxonomy, verify recipe and failure table live in skills/bt-gauntlet/references/.
How a part is judged
- The bar is an image, and it is locked. No prose bars. If you have no reference media it generates one (a real in-engine screenshot, never concept art), then freezes it for the run.
- Evidence comes from locked cameras. Three to five named cameras with exact transforms; a frame shot from anywhere else is inadmissible.
- A scored rubric, not a vibe. Composition / Lighting & Atmosphere / Materials / Detail / Motion, out of 10 (models get their own axes). Default pass mark 8.0.
- The critic returns the whole gap list, prioritised and actionable — not one gap per round — and it sees the previous round's score and screenshot so regressions are caught.
- Hard gates are separate from the score: perf budget, console clean, camera lock, component authority, and the pipeline's own gate (e.g. did the change actually cross the Unity export boundary).
- A stall ladder forces adaptation. No full-point gain in two rounds, or the same top gap named twice → incremental tweaks are forbidden and one architectural change is required. If that fails too → it stops and asks you.
1. Start a new gauntlet (the normal way)
/bt-gauntlet I want you to build a first-person shooter at the level of the most
recent Call of Duty games. It should be utterly perfect, visually beautiful, with
every single thing done at AAA quality—from textures to physics to anything you
could think of.What happens:
- It derives a job name from the brief (e.g.
cod-fps) and runs the interview — deliverable kind and deliverable, objective, the target image and locked cameras (attach screenshots, or let it generate the target), rubric and pass threshold, success condition, boundaries, kind-specific questions, loop mechanics. - It shows you the filled loop card and waits for your confirmation.
- It runs up to 5 rounds (the default session cap), then parks with a status report and the exact resume command.
Name it yourself and pick the template explicitly:
/bt-gauntlet --name:cod-fps --template:gauntlet I want you to build a first-person shooter ...Attach reference screenshots/clips with the message — they become the locked target in _gauntlet/cod-fps/reference/. If you attach nothing, the interview generates the target instead and shows it to you before anything is built.
2. Choose a template
| | --template:gauntlet (default) | --template:bounded |
| --- | --- | --- |
| Style | Full Shumer-style: decompose into parts, fan out builders, fresh harsh critic per part, scored against the locked target | Single-track loop card: one coherent improvement per round against an objective/metric/boundary checklist |
| Best for | One ambitious visual artifact chasing a reference bar ("CoD-level FPS") | Reliability- and cost-sensitive work with objective verifiers ("60 FPS, zero console errors, all checks green") |
| Cost profile | Heavier (builder + critic subagents per part) | Lighter (one improvement, one verifier per round) |
Both templates live verbatim in the SKILL.md, so you can edit their wording there.
3. Control how much one session does
/bt-gauntlet --rounds:20 --name:cod-fps <brief> # a long night: up to 20 rounds
/bt-gauntlet --rounds:1 --resume cod-fps # one careful round, then park
/bt-gauntlet --resume cod-fps # default: up to 5 rounds--rounds:N caps THIS invocation only. State is persisted after every round regardless, so even a hard cutoff (usage limit, closed laptop) loses at most the round in flight.
4. Stop, check things out, come back later (the resume cycle)
This is the core workflow the skill was built for:
# Tuesday night — start, run 5 rounds, it parks itself
/bt-gauntlet --name:cod-fps I want you to build a first-person shooter ...
# (any time) — peek without running anything
/bt-gauntlet --status cod-fps
# ...daily limit resets, or a week later, brand-new session, zero memory of the chat:
/bt-gauntlet --resume cod-fps--resume works because nothing lives in the conversation: the fresh session re-reads the Agent Reference, then loop-card.md → progress.md → the last round journals, sanity-checks them against the real project, reports "resuming from round N", and executes the recorded NEXT ACTION. Budgets count in rounds/attempts (cumulative across sessions), never wall-clock, so a two-week gap changes nothing.
Park a job deliberately mid-session:
/bt-gauntlet --stop cod-fpsYou can also just interrupt at any time — parking is only the polite version.
5. Run several gauntlets side by side
Each job owns its own _gauntlet/<name>/ folder — as many as you like, fully independent budgets, rounds, references, and evidence:
/bt-gauntlet --name:cod-fps <FPS brief>
/bt-gauntlet --name:racing-demo <racing brief>
/bt-gauntlet --status # index of ALL jobs, one line each
/bt-gauntlet --resume racing-demo # continue just that one
/bt-gauntlet --resume # only ONE job exists → resumes it;
# several exist → lists them and asksA new gauntlet whose name collides with an existing job is an error — it will offer --resume <name> or a different name, never overwrite.
6. Read the workspace while it's parked
Everything is plain markdown/media under _gauntlet/<name>/ — inspect it in your editor between sessions:
| File | What you'll find |
| --- | --- |
| loop-card.md | The confirmed template: objective, benchmark, success condition, boundaries. The loop never edits it. |
| brief.md | Your full interview answers. |
| progress.md | The resume brain: part checklist (- [ ] / - [~] / - [x]), attempt counts, failed-approaches log, budgets spent, and the one-line NEXT ACTION. |
| rounds/round-NN.md | Per-round journal: what was built, the critic's verdict and the largest gap it named, evidence links. |
| reference/ | Your benchmark screenshots/clips — what critics compare against, blind. |
| evidence/ | Captured browser screenshots + perf numbers from our side of the A/B. |
Want to redirect the loop? Edit progress.md's NEXT ACTION (or part priorities) before resuming — the loop trusts the files.
7. Skip the interview (--card: — non-interactive start)
Hand it a pre-filled loop card (either template with every slot answered, including <NAME>) and it starts immediately, no questions:
/bt-gauntlet --card:_specs/racing-game_gauntlet-card.md --rounds:10If any slot is unfilled or vague it STOPS and tells you which — it never guesses. This is the entry point used by the spec workflow below, and handy any time you want to author the card by hand.
8. Use it with the spec workflow (bt-spec / bt-plan / bt-execute)
The gauntlet contains its own plan — never run bt-plan on gauntlet work. Its loop card is the spec, its part checklist in progress.md is the plan, its rounds are the execution.
Rule of thumb:
- Acceptance criteria you can enumerate up front, each done in one pass → bt-spec → bt-plan → bt-execute.
- "As good as that reference", unknown number of iterations → bt-gauntlet.
Typical two-phase sequence for a real game:
# Phase 1 — foundation (checklist-shaped work: scaffold, controller, physics, HUD)
/bt-spec <feature brief>
/bt-plan @_specs/<feature>_spec.md
/bt-execute @_specs/<feature>_plan.md ALL
# Phase 2 — quality (reference-bar work on top of what now exists)
/bt-gauntlet --name:aaa-polish <polish brief + reference media>Or fold the gauntlet INTO a spec run by naming it in the brief:
/bt-spec Build the racing game, then polish it to Gran Turismo quality using bt-gauntletIn that flow, bt-spec runs the gauntlet interview at spec time and writes the pre-filled card (_specs/<feature>_gauntlet-card.md); bt-plan makes the gauntlet the final task invoking --card:; bt-execute re-enters a parked job via --resume <name> on each NEXT and flips the task's checkbox only when the job genuinely reports DONE.
9. Which method should I use for making games? (recommendation)
Recommendation: the two-phase hybrid — spec/plan/execute for the foundation, then a gauntlet job for the quality bar. One-shot gauntlet looping is the demo; the hybrid is how you'd actually ship.
Here's the reasoning:
Why not one-shot gauntlet from a cold start (the pure Claude-of-Duty move): it makes the loop do work loops are bad at. Scaffolding, the player controller, physics wiring, input, level loading — these have crisp, enumerable acceptance ("Havok is initialized, capsule controller walks the test level at 60 FPS"). Running builder-vs-critic rounds on that is paying five critics to confirm a checkbox. It's also where the loop's weaknesses bite: the decomposition emerges mid-run instead of being engineered, coupled foundation systems tempt bad fan-out, and — worth remembering — even Shumer's original run never won a single blind A/B against real CoD. The honest lesson of that experiment is that the demanding reference kept the agent improving; not that one prompt replaces engineering.
Why not spec/plan/execute alone: it terminates at "acceptance met," which for visuals means "good for AI." There's no mechanism that keeps pushing after the box is checked — no reference comparison, no fresh critic naming the largest gap, no "go again." That's precisely the gap the gauntlet fills.
So the playbook for a real game:
# Phase 1 — foundation (cheap, deterministic, checkbox-resumable)
/bt-spec <game systems brief>
/bt-plan @_specs/<feature>_spec.md
/bt-execute @_specs/<feature>_plan.md ALL
# Phase 2 — quality (reference-chasing, critic-judged, round-resumable)
/bt-gauntlet --name:aaa-polish <polish brief + reference screenshots>Two refinements on top:
- Scope each gauntlet job tightly. Rather than one giant
aaa-polishjob, consider a couple of focused jobs — e.g.visuals(lighting, materials, post-processing vs reference frames) andgame-feel(movement, weapon feedback, hit reactions). Focused jobs give critics sharper rubrics, keep parts genuinely independent for fan-out, and let you park/resume/abandon them separately. - Reserve true one-shot gauntlet (
/bt-gauntlet <ambitious brief>on an empty folder) for what it's genuinely great at: prototypes, jams, and "show me what's possible" experiments where you want the decomposition to emerge and the ride is the point. That's the mode the Call-of-Duty example lives in, and it works — it's just not the cost-efficient path to a shippable game.
The interplay is already wired for this: phase 2's interview can point its reference media and priorities at exactly what phase 1 built, and if you want it fully hands-off you can fold phase 2 into the spec run itself (/bt-spec … then polish to <reference> quality using bt-gauntlet), which pre-fills the loop card and drives it via --card: — per section 8 above.
10. Unattended runs (optional native loop drivers)
The gauntlet needs no native loop tool — the round protocol is the loop, and --resume is the continuation. But because state persists after every round, you can safely layer a host's re-invocation driver on top for hands-off runs:
# Claude Code — keep re-invoking resume rounds while you sleep
/loop /bt-gauntlet --resume cod-fps
# Hosts with a persistent goal feature (e.g. /goal) — same idea:
# keep "/bt-gauntlet --resume cod-fps until the job reports DONE" alive across turnsWhy this is safe: whenever the native driver dies — daily/weekly limit, session end, closed laptop — the job is simply parked, exactly as if you had stopped it yourself. Nothing is lost beyond the round in flight, and /bt-gauntlet --resume cod-fps in a fresh session continues where it left off.
Keep the layering straight:
| | The native driver (/loop, /goal) | The gauntlet |
| --- | --- | --- |
| Job | When to re-poke the agent | What a round is: build → evidence → fresh critic → record |
| Survives session end / limits | No | Yes — _gauntlet/<name>/ + --resume |
| Portable across hosts | Each host different; some have nothing | Identical everywhere |
This is an operator convenience you apply from the outside. The SKILL.md deliberately never references any native loop tool, so the skill stays fully portable — on a host with no such feature, you just re-invoke --resume yourself.
11. How a gauntlet ends
In order of precedence, a round's gate check ends things when:
- Success condition met → one final fresh integration critic inspects the whole game for seams and consistency → report DONE with the evidence summary.
- The stall ladder reaches STALLED — an architectural change was forced after two flat rounds (or the same top gap twice) and still did not move the score → park + escalate, rather than burn rounds guessing at another big swing.
- A loop-card boundary fires (total rounds exhausted, attempts-per-part hit without a new strategy, repeated blocker, permission needed) → park + escalate to you with specifics.
- The
--roundssession cap is reached → park cleanly and print/bt-gauntlet --resume <name>.
Deploy, spending, credentials, deletion, and messaging are always behind your explicit approval, no matter what the brief says.
Quick reference
| Command | Effect |
| --- | --- |
| /bt-gauntlet <brief> | New gauntlet: interview → confirm card → loop (≤5 rounds) → park |
| /bt-gauntlet --name:x --template:bounded --rounds:10 <brief> | New named job, bounded template, 10-round session |
| /bt-gauntlet --card:<file> [--rounds:N] | New job from a pre-filled card, zero questions |
| /bt-gauntlet --resume [x] | Continue a job in any (brand-new) session |
| /bt-gauntlet --status [x] | Read-only: one job's detail, or the index of all jobs |
| /bt-gauntlet --stop [x] | Park a job deliberately, print its resume command |
Babylon Toolkit Agent Persona (@babylonjs-toolkit/agent)
Simply ask your agent to install the Babylon Toolkit Agent for you.
