badgr-cli
v1.2.11
Published
Badgr — run or serve GPU workloads from one command
Maintainers
Keywords
Readme
badgr-cli
Badgr supports many GPU workloads through two commands: serve for persistent endpoints, run for jobs.
Safety promise: every run has
--max-cost, live logs, automatic teardown, and a receipt. Runbadgr down <id>any time to stop billing immediately.
npm install -g badgr-cliJump to: Quick start · Coding agents (badgr launch) · Image generation · badgr doctor · badgr diagnose · Commands · serve options · run options · Receipts · OpenAI compatibility · GPU options · Advanced · Requirements
Quick start
# 1. Authenticate once
badgr login
# 2. Verify the stack end-to-end
badgr test
# 3. Run a project folder on a GPU
badgr run . --cmd "python train.py" --max-cost 5
# 4. Serve an OpenAI-compatible inference endpoint
badgr serve meta-llama/Llama-3.1-8B-Instruct --max-cost 10
# 5. Stop billing
badgr down <deployment-id>
# 6. View cost, route, and retry receipts
badgr receiptsbadgr serve prints a URL you can point any OpenAI SDK client at — see OpenAI compatibility.
Coding & testing agents: badgr launch
Run a coding agent or a test suite on a CPU VM with one command — no image, source, or --max-cost required for the four built-in workloads. Three separate things are going on here, worth keeping straight: which agent CLI runs (cline, claude, codex, playwright), who pays for model usage (Badgr, or your own Anthropic/OpenAI account), and who provisions the VM and runs the command (always Badgr).
badgr launch cline "Fix the checkout bug" # Badgr provides model access — no account to connect
badgr launch claude "Fix the checkout bug" # runs Claude Code — connect your Anthropic account
badgr launch codex "Write tests" # runs the Codex CLI — connect your OpenAI account
badgr launch playwright "Test the checkout flow" # no model account involved at allAuthentication model:
| Workload | CLI that runs | Who pays for model usage |
|----------|---------------|---------------------------|
| cline | Cline | Badgr provides and pays for model access |
| claude | Claude Code | Your Anthropic account — Badgr only provides the VM |
| codex | Codex CLI | Your OpenAI/ChatGPT account — Badgr only provides the VM |
| playwright | Playwright | No model account required |
Installing/running the claude or codex CLI doesn't by itself give a disposable Badgr VM model access — the CLI still needs to authenticate, and a VM's sign-in doesn't inherit from your laptop. badgr launch claude/badgr launch codex prompt inline the first time to connect your account and store it for next time; cline and playwright need nothing. Each workload picks its own VM size automatically (small, or browser for playwright's preinstalled Chromium) — override with --size small|medium|browser. A $2 default --max-cost applies unless you pass your own.
Today, "connect your account" means securely storing an API key (ANTHROPIC_API_KEY / OPENAI_API_KEY) — both Claude Code and the Codex CLI also support signing in via a Claude.ai/ChatGPT account, and a future version of badgr connect may add that OAuth flow instead of a pasted key; the one-time-connection UX is the same either way.
badgr connect anthropic --key sk-ant-... # connect your Anthropic account ahead of time (optional — badgr launch prompts inline if missing)
badgr connect # list what's already connected
badgr connect anthropic --remove # disconnectRetrieve results after a run finishes:
badgr pull <deployment-id> # pull a code-editing agent's patch as a local git diff/branch
badgr artifacts <deployment-id> # download everything else — test reports, screenshots, tracesBare-task form — no agent name, no repo flags
For a quick "go fix this" against a repo you're already sitting in, skip the agent name and let Badgr resolve everything else:
badgr launch "Fix the checkout bug"This resolves a repo (explicit --repo → the current local git repo → the last repo a bare-task launch verified against), infers an acceptance/"done" command from the repo's own evidence (package.json scripts.test, pytest, a Makefile target, or CI config — never a guessed shell command), and runs it as the verification gate by default instead of trusting the agent's own "done" claim. Result is reported as ✓ Verified (with the commit/PR it produced) or NOT VERIFIED — never inferred from the agent's own text.
badgr launch "Fix the checkout bug" --repo https://github.com/org/repo # explicit repo, skips resolution
badgr launch "Fix the checkout bug" --check "npm test" # explicit acceptance command, skips inferenceIf no repo can be resolved, or no acceptance command can be inferred, launch fails and asks for --repo/--check rather than launching unverified. A repo only gets remembered for next time once a launch against it actually verifies.
Advanced escape hatch — any other command
For anything beyond the four built-in workloads, run an arbitrary command on a CPU VM the same way badgr run does for GPU jobs:
badgr launch . --max-cost 1 -- npm test
badgr launch https://github.com/user/repo --max-cost 1 -- python narrgo.pyEverything before -- is a badgr launch flag; everything after -- is passed to your command verbatim, including anything that looks like a flag (badgr launch . -- claude --max-cost 1 sends --max-cost 1 to claude, not to Badgr). The default runner image has no Node.js/npm — use --image <image> for anything that needs it.
| Flag | Default | Description |
|------|---------|-------------|
| --cmd "<command>" | — | Quoted command form, equivalent to -- <command> |
| --image <image> | CPU runtime default | Custom image instead of the default |
| --detach / --no-detach | detach | launch detaches by default; --no-detach streams logs and waits |
| --env KEY=VALUE | — | Environment variable (repeatable) — Badgr warns if a value looks like a secret |
| --artifacts <path> | — | Extra path to capture and upload (repeatable), e.g. --artifacts playwright-report. Retrieve with badgr artifacts <id> |
| --max-cost <$> | $2 | Auto-stop when spend reaches this amount |
| --max-runtime <min> | 60 | Auto-stop after N minutes |
| --region US\|EU\|AU | — | Region preference |
| --size small\|medium\|browser | per-workload default | VM class override |
badgr job — a tracked coding-agent job
badgr job cline "Run the Chromium tests and tell me what failed" --check "npm run test:chromium" --max-cost 1Submits via POST /v1/jobs (type: "agent") — the same Jobs API used by every other job type. Requires --check <command> (verifies success); tracked at /jobs with a job_id, status, logs, output, cost, and time.
Also try: image generation
badgr comfyui batch --workflow sdxl-basic --prompt "a cat on a beach" --max-cost 5Runs a blessed ComfyUI workflow, no setup, and prints image URLs when done. No manual teardown needed — it stops itself. Details in badgr comfyui batch options.
Something not working? badgr doctor
Read-only diagnosis for GPU workload failures — checks your GPU, drivers, CUDA, PyTorch, and disk, then tells you what's likely wrong and what to try next. No login, no setup, never mutates your machine.
badgr doctor# A model won't fit / you're not sure it'll fit before you try
badgr doctor --model meta-llama/Llama-3.1-8B-Instruct
# You have an error from a crashed job — save it to a file first
badgr doctor --logs error.log
# A ComfyUI workflow is failing or referencing something missing
badgr doctor --workflow my-workflow.json
# You started a server and it's not responding (ComfyUI, llama.cpp, and
# generic health formats are recognized too, not just OpenAI-style)
badgr doctor --url http://localhost:8000/v1/models
# Splitting a big model across multiple GPUs (tensor parallel)
badgr doctor --model meta-llama/Llama-3.1-70B-Instruct-AWQ --serve --gpu-count 2
# Machine-readable output for scripts/CI
badgr doctor --jsonRun badgr doctor --help for the full flag list. Details in
docs/gpu-doctor.md.
Badgr Smoke Test: badgr diagnose
Paste a GitHub issue, a Docker image, a repo URL, a log file, a ComfyUI
workflow, or a raw conversation — Badgr auto-detects the input, redacts
secrets, and runs free static checks as part of a Badgr Smoke Test.
Nothing runs on a GPU without explicit --approve.
# Free Badgr Smoke Test — no login required
badgr diagnose "https://github.com/org/repo/issues/123"
badgr diagnose ajayrajtp/vllm_gemma412b:latest
badgr diagnose ./vllm-error.log
badgr diagnose "https://github.com/org/repo"
badgr diagnose workflow.json
# Free mechanical validation of the produced command — still no GPU, no login
badgr diagnose "https://github.com/org/repo/issues/123" --smoke
# Approve a capped smoke test after diagnosis (login required)
badgr diagnose "https://github.com/org/repo/issues/123" --approve
# Resume an existing case, e.g. one shared via a case link
badgr diagnose repro_xxxxxxxx --approve
badgr diagnose "https://aibadgr.com/repro/repro_xxxxxxxx" --approve
# Machine-readable output
badgr diagnose "https://github.com/org/repo/issues/123" --json| Flag | Description |
|------|-------------|
| --smoke | Free. Runs real, cheap mechanical validation of the produced command (syntax, CLI-entrypoint corroboration, referenced-file presence, Docker ENTRYPOINT/CMD consistency, required env vars, plus a client-side check of any local file path you pasted). Never starts a GPU, never requires login. Combine with --approve — the smoke checks print first, then the normal approve flow runs. Not the same as badgr run --smoke, which launches a real, billable GPU job |
| --approve | Approve the capped smoke test after diagnosis (opens browser sign-in automatically if not logged in) |
| --docker <image> | Force Docker-image intake (override auto-detect) |
| --repo <url> | Force repository intake (override auto-detect) |
| --comfyui <path> | Force ComfyUI workflow intake (override auto-detect) |
| --github <url> | Include a GitHub issue URL found inside pasted text as additional context (opt-in, never fetched automatically) |
| --json | Machine-readable JSON output |
Every run prints one status: NEEDS INFO (fields still missing) →
READY (complete evidence-backed command, no smoke run) → SMOKE CHECKED
(--smoke ran, every applicable check passed or was skipped) / INVALID
(--smoke ran and a check actually failed) → VERIFIED (reserved for an
actual successful GPU-provisioned run — never assigned from static
resolution or --smoke). READY/SMOKE CHECKED results also print the
canonical badgr run/badgr serve command --approve would run, plus a
shareable case link — pass that case ID or URL back into badgr diagnose
<case_id_or_url> --approve to resume the same case without re-diagnosing.
badgr run-issue is an alias for badgr diagnose, matching the
aibadgr.com/run-issue Badgr Smoke Test
web flow.
Commands
login
connect
node
doctor
diagnose
run
launch
job
serve
status
logs
pull
artifacts
down
restart
rerun
heartbeat
receipts
test| Command | What it does |
|---------|-------------|
| badgr login | Save API key to ~/.badgr/config.json |
| badgr connect <provider> | Store a provider credential (anthropic, openai, openrouter, deepseek, glm, custom) for badgr launch |
| badgr node connect\|list\|inspect\|disable\|remove | BYO GPU — connect your own Linux/NVIDIA or Linux/AMD-ROCm machine and run run/serve on it (see BYO GPU) |
| badgr doctor | Diagnose a GPU workload failure — read-only, no login needed |
| badgr diagnose "<input>" | Run a Badgr Smoke Test on a GitHub issue, Docker image, repo, log, or workflow — free, no GPU until --approve |
| badgr run <command> | Run a one-off GPU job (any container command) |
| badgr launch cline\|claude\|codex\|playwright "<task>" | Run a coding/testing agent on a CPU VM — image + command auto-selected |
| badgr launch "<task>" | Bare-task form — resolves a repo and acceptance command automatically, verifies the result |
| badgr job <agent> "<instruction>" --check "<cmd>" | Tracked coding-agent job via POST /v1/jobs (type: agent) |
| badgr serve <model> | Start a persistent OpenAI-compatible endpoint |
| badgr status | Show what's running and what's billing |
| badgr logs <id> | Fetch log output from a deployment |
| badgr pull <id> | Pull a code-editing agent's patch as a local git diff/branch |
| badgr artifacts <id> | Download non-patch outputs (test reports, screenshots, traces) |
| badgr down <id> | Terminate a deployment — stops billing immediately |
| badgr restart <id> | Relaunch an endpoint with the same config, on a new deployment ID, keeping its API key |
| badgr rerun <id> | Replay a past job or endpoint with its exact original spec, on a new deployment ID |
| badgr heartbeat <id> | Reset an endpoint's idle-timeout clock (see --idle-timeout under badgr serve options) |
| badgr receipts [n] | Cost, route, and retry receipts (default 10) |
| badgr test | Run an end-to-end test (provision → run → teardown) |
More commands below, under Advanced: comfyui, train, transcribe, embed, workload, workspace, batch, sbatch, capacity, billing, models, template.
badgr serve — for anything that needs a persistent endpoint: LLM serving, embeddings, image generation APIs, transcription APIs.
badgr run — for anything that starts, runs, and exits: batch inference, fine-tuning, evals, image/video batch jobs, audio processing.
Common flags
These appear on most commands (run, serve, comfyui, train, transcribe, embed) — documented once here instead of repeated in every table below.
| Flag | Default | Description |
|------|---------|-------------|
| --gpu <type> | auto | GPU type override — see GPU options |
| --target <node> | — | Run on one of your own connected machines instead of Badgr's cloud capacity — a node name or node_… id from badgr node list. See BYO GPU |
| --tier 1\|2 | 1 | 1 = reliable managed routing (default); 2 = lower-cost marketplace routing |
| --region US\|EU\|AU | — | Region preference. If omitted, Badgr chooses best available capacity |
| --max-price <$/hr> | — | Hard spend cap per GPU-hour |
| --count <n> | 1 | Number of GPUs |
| --env KEY=VALUE | — | Environment variable (repeatable) |
| --max-cost <$> | computed | Auto-stop when total spend reaches this amount — optional for run and serve (a conservative default is computed from the GPU tier and runtime if omitted); required (or --persistent) for comfyui run, comfyui batch, train lora |
| --detach | off | Launch and return immediately, don't stream logs (serve uses --no-wait instead — see below) |
Each command section below lists only its own extra flags.
badgr serve options
badgr serve meta-llama/Llama-3.1-8B-Instruct --gpu L40S --region EUEndpoints bill continuously, so serve always runs under a spend cap: pass --max-cost explicitly, use --persistent to run until manually stopped, or omit --max-cost and Badgr computes a conservative default from the GPU tier and runtime (printed before provisioning, with a note that it's a default — override it with your own --max-cost). --dry-run previews the plan without provisioning.
| Flag | Default | Description |
|------|---------|-------------|
| --image <img> | — | Serve a custom container instead of a HuggingFace model |
| --task <task> | — | vLLM task override, e.g. embed for embedding models |
| --runtime llama.cpp\|ollama | vLLM | Serve via a different runtime instead of vLLM |
| --hf-repo <repo> | — | HuggingFace repo for a GGUF file (with --runtime llama.cpp), e.g. org/model-repo |
| --hf-file <file> | — | GGUF filename within that repo, e.g. model.gguf (required with --runtime llama.cpp) |
| --idle-timeout <min> | — | Auto-stop after N minutes with no badgr heartbeat call — see badgr heartbeat |
| --persistent | off | Run until manually stopped — satisfies the cost-control requirement in place of --max-cost |
| --check-nodes <n1,n2> | — | For ComfyUI-shaped images: verify custom nodes are installed after startup |
| --health-path <path> | auto | Readiness path to poll (auto-detected: ComfyUI → /system_stats, llama.cpp → /health) |
| --no-wait | off | Skip endpoint health check and return immediately |
| --yes / -y | off | Skip duplicate-deployment warning |
| --dry-run | — | Preview the plan (GPU, price) without provisioning |
| --list-aliases | — | List blessed vLLM model aliases (qwen-7b, llama-8b, qwen-coder-7b) and exit — no provisioning, no API key required |
Blessed aliases expand to a full model ID + preset GPU, e.g. badgr serve qwen-7b → Qwen/Qwen2.5-7B-Instruct on an RTX 4090. Run badgr serve --list-aliases to see the current list.
# Serve a Hugging Face GGUF file via llama.cpp instead of vLLM
badgr serve --runtime llama.cpp --hf-repo org/model-repo --hf-file model.gguf --max-cost 10
# Serve Open WebUI, a chat UI, pointed at a model endpoint
badgr serve openwebui --model qwen-7b --max-cost 10
badgr serve openwebui --connect <existing-endpoint-url> # connect to an endpoint you already have
# Auto-stop only when idle — requires periodic badgr heartbeat calls to stay up
badgr serve qwen-7b --idle-timeout 30 --max-cost 10Model support levels
badgr serve qwen-7b is the happy path — a tested route with no extra setup. badgr serve also accepts any other model ID or a custom container:
| Level | What it means |
|-------|---------------|
| Tested route (badgr serve qwen-7b) | One of the aliases above — tested and officially supported. |
| Best-effort Hugging Face model (badgr serve <org>/<model>) | Any other Hugging Face model ID. Badgr will try a compatible route — not a guarantee every model works. |
| Custom container (badgr serve --image ...) | You own the server behavior; Badgr manages runtime, logs, spend caps, teardown, and the receipt. |
Gated Hugging Face models (e.g. Llama, Gemma) may need --env HF_TOKEN=$HF_TOKEN. Badgr only prints that hint if the deployment actually fails to start.
BYO serve from an existing script
If you already have a working GPU launch script, this is the whole flow:
badgr node connect --name gpu-box-1
badgr serve ./start-vllm.sh --target gpu-box-1No --mount, --image, --cmd, --port, --health-path, or --verify-model needed. Badgr reads the script (byte-for-byte — never reformatted, re-indented, or rewritten), infers what it needs from the script's own evidence, prints a short resolved plan, and launches:
Preflight
✓ AMD workload
✓ gfx1030
✓ 4 GPU(s) required
✓ Port 8000
✓ Health: /v1/models
✓ Model: linuxai-01/Qwen3.8-27B
✓ 5 host path(s) found: /home/faisal/.cache, /home/faisal/venvs/vllm-rocm-0.28.0, ...
✓ Runtime image: rocm/dev-ubuntu-22.04:6.1-complete (inferred — not a compatibility guarantee for your own venv/build)
Starting endpoint...What it infers from the script, and how:
| Field | Evidence |
|-------|----------|
| GPU vendor (AMD/NVIDIA) | ROCm/CUDA env markers (PYTORCH_ROCM_ARCH, CUDA_VISIBLE_DEVICES, ...) or a gfx####/rocm token — same rules BYO Preflight already uses |
| GPU architecture | PYTORCH_ROCM_ARCH, HSA_OVERRIDE_GFX_VERSION, a bare gfx####, or TORCH_CUDA_ARCH_LIST |
| GPU count | --tensor-parallel-size N, or a comma-separated *_VISIBLE_DEVICES list |
| Port | --port N or a PORT= assignment |
| Health path | /v1/models for a vllm serve script; otherwise not assumed |
| Served model (for real inference verification) | --served-model-name, else the model path right after vllm serve |
| Host mounts | Absolute paths in source .../activate, KEY=/path env assignments, and --flag /path args — a venv's bin/activate widens to the venv root (needed to actually run it), a printf-style save pattern (results%d.csv) widens to its directory, sibling paths under a shared .cache root collapse into one mount. System/credential paths (/etc, /proc, /sys, /dev, ~/.ssh, /, ...) are never auto-mounted. Each detected path is checked against the connected node's own filesystem (its worker daemon answers the check — not just this CLI process's local disk, which isn't guaranteed to be the same machine) — if one is missing, Badgr stops with NEEDS INFO naming it, before starting anything. Falls back to a local-only check (with a printed warning) if the node can't be reached for the check |
| Runtime image | A best-effort ROCm or CUDA base image by detected vendor — not a compatibility guarantee; your own venv/build still needs to actually run inside it. If vendor can't be determined, Badgr stops with NEEDS INFO asking for --image rather than guessing. If BYO automatic recovery has to swap to a different working image, Badgr says so (Image (corrected):) instead of pretending the original guess is what ran — and that corrected image is what a repeat run of this same script on this node starts from next time, not the shallow default again |
Every inferred field is overridable with the same explicit flags raw-command mode already takes — an explicit flag always wins over detection:
badgr serve ./start-vllm.sh --target gpu-box-1 \
--image my/pinned-rocm:tag --port 8000 --health-path /v1/models --verify-model my-model \
--mount /extra/path:/extra/pathPrecedence: explicit flag → script evidence → safe default → NEEDS INFO. Badgr never guesses past NEEDS INFO when a wrong guess could make the workload invalid (an unresolvable image, a missing host path) — it names the one unresolved thing rather than failing generically.
BYO raw-command endpoint mode (advanced)
Script mode above is really a friendlier front end for these same primitives — reach for them directly when you'd rather build the command/mounts yourself, or don't have a script file (e.g. composing it inline):
badgr serve --target gpu-box-1 --image rocm/vllm:latest \
--cmd "source /home/you/venvs/vllm-rocm/bin/activate && PYTORCH_ROCM_ARCH=gfx1030 vllm serve /models/my-model --host 0.0.0.0 --port 8000" \
--port 8000 \
--persistent--cmd requires --target (your own connected node — see BYO GPU and --mount for making host paths like a venv or model directory visible inside the container), --image, and --port (Badgr can't infer a port from an opaque command). It's mutually exclusive with a model — the command already says what to load. --health-path defaults to /v1/models (vLLM's own convention) if not given; pass an explicit one for a non-vLLM server. Badgr provisions, tracks, health-checks, and gives you logs/restart/receipts for this exact command — it never rebuilds or reinterprets it, including on automatic BYO recovery retries (only the image can be swapped there, never the command).
Badgr also proves real inference, not just that the port opened: it auto-detects a servable model id from --served-model-name or a vllm serve <model> invocation inside --cmd and sends it a real POST /v1/completions once the endpoint reports ready — the same check a normal badgr serve <model> gets. If nothing in --cmd looks like a vllm serve command, Badgr says so up front (Verify: readiness/port-health only) rather than silently only checking that the port is open; pass --verify-model <name> to verify inference for a non-vLLM server or a name detection got wrong.
--mount HOST:CONTAINER[:ro] (repeatable) on badgr serve applies to this job specifically, in addition to whatever the node was connected with (badgr node connect --mount) — this is what script mode uses under the hood to hand over the host paths it found without requiring them pre-configured at connect time.
This mode is restricted to --target (your own connected node) — a cloud-provider pod running an arbitrary caller-supplied command is a materially larger trust boundary and is out of scope.
badgr run options
Three source patterns:
# Flow 1 — local project folder (primary)
badgr run . --cmd "python train.py" --max-cost 5
# Flow 2 — public GitHub repo
badgr run https://github.com/user/repo --cmd "python train.py" --max-cost 5
# Flow 3 — custom Docker image (advanced)
badgr run . --image mycompany/custom:latest --cmd "python train.py" --max-cost 5--cmd is optional for local-folder and GitHub-repo sources — if omitted, Badgr looks for a recognizable entrypoint file (e.g. train.py, main.py, app.py) and a Dockerfile, the same detection either source uses (a GitHub repo is read via one contents-listing API call, never cloned). Low-confidence detection still requires an explicit --cmd. Pass --cmd yourself any time to override.
Badgr zips and uploads the folder (Flow 1) or clones the repo (Flow 2), picks a generic runner, installs deps, runs the command, stores outputs for 48 hours, and tears down the GPU. --max-cost is optional — omit it and Badgr computes a conservative default from the GPU tier and runtime; pass your own to override.
# No GPU needed — run on a plain CPU VM instead
badgr run . --cmd "npm test" --no-gpu --max-cost 1
# Describe basic compute needs instead of a GPU model — Badgr finds a compatible machine
badgr run . --cpu 16 --memory 64GB --gpu-memory 24GB --max-cost 5| Flag | Default | Description |
|------|---------|-------------|
| --cmd <command> | inferred | Command to run inside the uploaded project or cloned repo — inferred from an entrypoint file/Dockerfile for folder/GitHub flows if omitted (falls back to requiring --cmd on low-confidence detection); required for the custom-image flow |
| --min-vram <GB> | — | Minimum VRAM in GB — optional constraint for Auto routing (alias: --gpu-memory) |
| --cpu <cores> | — | Minimum CPU cores (for a CPU-only run) |
| --memory <size> | — | Minimum RAM, e.g. 64GB |
| --no-gpu | off | Run on a CPU-only VM — no GPU is provisioned (conflicts with --gpu/--min-vram) |
| --image <img> | — | Custom Docker image — bypasses the runner |
| --max-runtime <min> | — | Auto-stop after N minutes |
| --save <name> | — | Save this job as a named workload after it completes |
| --workspace <name-or-id> | — | Link this run to a workspace (name or ws_ ID) |
| --output <path> | — | See "Crash recovery" below |
| --checkpoint <path> | — | See "Crash recovery" below |
| --retry-safe | off | See "Crash recovery" below |
| --resume-cmd "<cmd>" | — | See "Crash recovery" below |
Crash recovery is a convention, not a feature Badgr runs for you. Badgr just passes these through into the container as environment variables — your script has to read them and actually write the files:
| Flag | Container sees | Your script needs to |
|------|-----------------|-----------------------|
| --output <path> | BADGR_OUTPUT_DIR=<path> | Write outputs it wants preserved to that path |
| --checkpoint <path> | BADGR_CHECKPOINT_DIR=<path> | Write/read resumable checkpoints at that path |
| --retry-safe | BADGR_RETRY_SAFE=1 | Confirm it's safe to re-run from scratch (e.g. it checkpoints/dedupes internally) |
| --resume-cmd "<cmd>" | (not passed to the container) | Nothing — Badgr just prints this command back to you on failure, so you have it without digging through shell history |
On failure, badgr run / badgr serve / badgr comfyui run all print a Class:/Next: pair identifying what went wrong (e.g. out_of_memory, image_pull_failed, cuda_unavailable) and a suggested next step.
badgr comfyui run options
badgr comfyui run workflow.json --max-cost 10
badgr comfyui run workflow.json --gpu RTX_4090 --check-nodes KSampler,CLIPTextEncode
badgr comfyui run workflow.json --max-cost 10 --comfyui-args "--disable-dynamic-vram"Requires either --max-cost or --persistent to prevent runaway billing.
| Flag | Default | Description |
|------|---------|-------------|
| --check-nodes <n1,n2> | — | Verify custom nodes are installed after startup |
| --comfyui-args "<extra ComfyUI flags>" | — | Extra ComfyUI CLI launch flags, e.g. --comfyui-args "--disable-dynamic-vram" or --comfyui-args "--disable-cuda-graphs --lowvram". Parsed with shell-compatible quoting into a structured argument list — Badgr handles translating that into whatever the current image's launch mechanism needs; you never need to know or set that yourself. Conflicts with --env for the same underlying variable (see below) |
| --no-wait | off | Skip health check, return immediately |
| --persistent | off | Run until manually stopped (no spending cap) |
| --yes / -y | off | Skip duplicate-deployment warning |
Before provisioning, Badgr prints what it resolved, e.g.:
ComfyUI args: --disable-dynamic-vram--comfyui-args is the supported interface — don't set the underlying runtime variable directly. As an advanced escape hatch, today's image reads CLI_ARGS and it remains available via --env CLI_ARGS=..., but combining the two is rejected rather than silently merged or one silently winning:
Conflicting ComfyUI launch arguments: use --comfyui-args or --env CLI_ARGS=..., not both.Swapping the underlying ComfyUI image later never changes this contract — only how Badgr translates --comfyui-args internally.
badgr comfyui batch options
Productized batch image generation — runs a list of prompts through a blessed ComfyUI workflow and returns image URLs. No ComfyUI setup, no workflow file, no manual teardown.
badgr comfyui batch --workflow sdxl-basic --prompts prompts.txt --max-cost 10
badgr comfyui batch --workflow sdxl-basic --prompt "a cat on a beach" --prompt "a dog in the park" --max-cost 5
badgr comfyui batch --workflow flux-basic --prompt "a neon city at night" --max-cost 5Blessed workflows: sdxl-basic (SDXL text-to-image, default sampler settings), flux-basic (FLUX.1-schnell text-to-image). Max 20 prompts per batch.
| Flag | Default | Description |
|------|---------|-------------|
| --workflow <name> | — | Blessed workflow ID (required) — sdxl-basic or flux-basic |
| --prompts <file> | — | Text file, one prompt per line |
| --prompt <text> | — | Inline prompt (repeatable) — combine with --prompts if needed |
| --max-runtime <min> | 60 | Auto-stop after N minutes |
| --gpu-type <type> | workflow default | GPU type override |
| --dry-run | — | Preview the batch (workflow, GPU, prompt count, cost) without provisioning |
Polls until complete and prints image URLs, or detaches with badgr status guidance if it outlives --max-runtime.
Receipts
Every badgr serve and badgr run action generates a receipt:
badgr receipts # last 10
badgr receipts 50 # last 50Each receipt includes runtime, estimated/settled cost, status, retries, teardown/billing result, and job/deployment ID.
Managing a deployment
badgr status # what's running and billing
badgr logs dep-abc123 # fetch current log output
badgr logs dep-abc123 --follow # stream logs until the deployment reaches a terminal state
badgr pull dep-abc123 # pull a coding agent's patch as a local git diff
badgr pull dep-abc123 --branch # ...as a new local branch instead of a diff
badgr pull dep-abc123 --diff-only # print the diff, don't touch the working tree
badgr artifacts dep-abc123 # download non-patch outputs (test reports, screenshots, traces)
badgr down dep-abc123 # terminate one deployment, stop billing
badgr down --all # terminate everything running, with a confirmation prompt
badgr down --all --yes # ...skip the confirmation prompt
badgr restart dep-abc123 # relaunch an endpoint with the same config — new ID, same API key
badgr rerun dep-abc123 # replay a past job/endpoint with its exact original spec — new ID
badgr heartbeat dep-abc123 # reset an endpoint's idle-timeout clock (see --idle-timeout on badgr serve)| Command | Flag | Description |
|---------|------|-------------|
| badgr logs <id> | --follow / -f | Stream/poll logs until the deployment reaches a terminal state instead of a one-shot fetch |
| badgr pull <id> | --branch | Create a new local git branch from the patch instead of leaving it as an unstaged diff |
| badgr pull <id> | --diff-only | Print the raw diff, don't touch the local working tree at all |
| badgr pull <id> | --yes / -y | Skip the confirmation prompt before applying |
| badgr down <id\|--all> | --all | Terminate every running deployment instead of one by ID |
| badgr down <id\|--all> | --yes / -y | Skip the confirmation prompt |
badgr restart is endpoint-only (it tears down the current pod and relaunches with the same GPU/model/price/runtime caps and endpoint API key, so existing clients keep working against a new URL). badgr rerun works for both one-off jobs and endpoints, replaying the exact original image/command/env/GPU/caps, and never tears down the source deployment.
OpenAI compatibility
badgr serve provisions a vLLM endpoint that is fully OpenAI-compatible:
import os
from openai import OpenAI
# Export BADGR_ENDPOINT from the URL printed by `badgr serve`
client = OpenAI(
api_key=os.environ["BADGR_API_KEY"],
base_url=os.environ["BADGR_ENDPOINT"],
)
resp = client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": "Hello"}],
)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.BADGR_API_KEY,
baseURL: process.env.BADGR_ENDPOINT, // URL printed by `badgr serve`
});GPU options
Badgr Auto selects the best eligible GPU for your workload. Add --gpu <type> or --min-vram <GB> only when you need more control.
| Flag value | GPU | VRAM | Best for | |-----------|-----|------|---------| | RTX_3090 | NVIDIA RTX 3090 | 24 GB | Dev, inference | | RTX_4090 | NVIDIA RTX 4090 | 24 GB | Inference, training, dev | | L40S | NVIDIA L40S | 48 GB | Inference, vLLM, embeddings | | A100 | NVIDIA A100 | 40–80 GB | Training, inference | | H100 | NVIDIA H100 | 80 GB | Large model training |
Additional GPU types may be routable depending on current capacity — check with badgr capacity. Pricing is confirmed before provisioning; use --dry-run to see it first. Full GPU support details: see NOTES.md.
BYO GPU: badgr node
Connect a machine you already own (Linux + NVIDIA or AMD/ROCm, Docker required) and run badgr run/badgr serve workloads on it instead of Badgr's cloud capacity — no separate execution path, just a --target on the commands you already use.
npm install -g badgr-cli # includes the node worker and its shared runtime
badgr node connect # register this machine and start the badgr-node worker
badgr node connect --name gpu-box-1 # optional human-readable name
badgr node list # your connected nodes, online/offline
badgr node inspect <node-id> # full JSON detail for one node
badgr node disable <node-id> # stop it from receiving new jobs (keeps the row)
badgr node remove <node-id> # disconnect and forget it
badgr run "python train.py" --target gpu-box-1 # run on that node instead of the cloud
badgr serve meta-llama/Llama-3.1-8B-Instruct --target gpu-box-1The published CLI package contains the worker's matching badgr-shared runtime. A fresh machine does not need a source checkout or a separate private package install before badgr node connect.
badgr node connect detects your GPU(s) via nvidia-smi (falling back to rocm-smi for AMD/ROCm) and Docker availability, registers the node (POST /v1/nodes), and — if a GPU, Docker, and python3 are all present — spawns the badgr-node worker daemon in the background so the node starts accepting jobs immediately. If anything's missing, it prints the exact command to start the worker manually once fixed. The worker itself re-probes vendor locally (nvidia-smi/rocm-smi) at startup and picks the matching docker run GPU flags — --gpus all on NVIDIA, or --device=/dev/kfd --device=/dev/dri --group-add video --group-add render --security-opt seccomp=unconfined on AMD.
A connected node is private by default — it is never marketplace-listed or shared; only you can target it, and only with --target. Marketplace listing, payout, and reputation scoring are not built yet (explicit follow-up phases).
Host-bound dependencies (venvs, patched source, model/cache dirs)
Jobs run inside a Docker container, which starts with none of this machine's filesystem visible except what's explicitly mounted. If your launch command depends on paths that already exist on this host — a Python venv, a patched source checkout (e.g. a custom-built vLLM), a model directory, a pip/HuggingFace cache — those paths will not exist inside the container unless you mount them:
badgr node connect \
--mount /home/you/venvs/vllm:/opt/venv \
--mount /home/you/vllm-src:/opt/vllm-src \
--mount /sync/Models:/models:ro \
--mount /home/you/.cache:/root/.cacheEach --mount is HOST:CONTAINER or HOST:CONTAINER:ro (same shape as docker run -v), repeatable, and both sides must be absolute paths. Mounts are configured once per node (not per job) and applied to every job that node runs. If you connected without --mount and the worker is already running, stop it and restart with the printed python3 .../node_worker.py ... command plus your --mount flags — there is no separate command to add a mount to an already-running worker yet.
An exact command that already works when you run it directly on this machine can still fail through --target if it references host paths you haven't mounted — that's a mounting gap, not a vendor/architecture/capacity one (see the BYO Preflight vendor/architecture checks above, which catch a different class of mismatch and pass independently of whether paths are mounted).
| Command | What it does |
|---------|-------------|
| badgr node connect [--name <name>] [--mount HOST:CONTAINER[:ro] ...] | Register this machine and start its worker |
| badgr node list | List your connected nodes |
| badgr node inspect <node-id> | Show one node's full detail as JSON |
| badgr node disable <node-id> | Stop a node from receiving new jobs |
| badgr node remove <node-id> | Disconnect and forget a node |
Routing
--tier 1 (default) uses managed provider routing — reliable, consistent performance. --tier 2 uses marketplace routing for lower-cost options. Most users should stick with the default.
Preview any command before provisioning with --dry-run, e.g.:
badgr serve meta-llama/Llama-3.1-8B-Instruct --dry-runAdvanced
Less common commands — training, transcription, embeddings, and the workload/workspace trackers. Same flags as run/serve unless noted (see Common flags).
badgr train / badgr train lora
badgr train config.yaml --gpu A100 --max-runtime 240 --env HF_TOKEN=$HF_TOKENDetects framework (axolotl, unsloth, trl) from the config file. Axolotl and TRL configs run today — unsloth/unrecognized configs are blocked before provisioning rather than billing a GPU that's guaranteed to fail. Default max-runtime is 120 min.
badgr train lora --base-model mistralai/Mistral-7B-v0.1 --dataset ./train.jsonl --preset small --max-cost 20
# Resume from a prior job's checkpoint instead of starting over
badgr train lora --base-model mistralai/Mistral-7B-v0.1 --dataset ./train.jsonl --resume https://.../checkpoint --max-cost 20Productized LoRA training — pass a base model and dataset, no Axolotl config file needed. Badgr generates the config from a preset and returns a downloadable adapter.
| Flag | Default | Description |
|------|---------|-------------|
| --framework <name> | auto-detect | Force framework: axolotl, unsloth, trl (axolotl/trl currently run; unsloth is blocked pre-provisioning) |
| --base-model <id> | — | HuggingFace model ID (required for train lora) — validated to exist before provisioning |
| --dataset <path\|url> | — | Local file, direct URL, or s3:// URI |
| --file-id <id> | — | Badgr upload ID instead of --dataset |
| --preset small\|medium | small | small = RTX 4090, rank 16, 3 epochs. medium = A100, rank 32, 5 epochs |
| --gpu-type <type> | preset default | GPU type override for train lora |
| --resume <checkpoint-url> | — | Continue training from a prior job's checkpoint instead of starting fresh |
| --dry-run | — | Preview the job without provisioning |
On completion, train lora prints an adapter_url — download with GET /v1/jobs/{job_id}/adapter, or via badgr workload info if saved.
badgr transcribe
badgr transcribe recording.mp3 --max-cost 2Whisper transcription. Accepts a public URL, S3/GCS URI, or a local file under 50 MB. Default max-runtime is 30 min.
| Flag | Default | Description |
|------|---------|-------------|
| --model <name> | large-v3 | Whisper model |
| --language <code> | — | Language hint (e.g. en, fr) |
| --output <format> | — | Output format: txt, srt, vtt |
badgr embed
badgr embed BAAI/bge-large-en-v1.5 documents.txt --max-cost 2Text embeddings via vLLM. Accepts a public URL, S3/GCS URI, or a local text file under 10 MB. Outputs JSONL ({"text": ..., "embedding": [...]}). Default max-runtime is 30 min.
| Flag | Default | Description |
|------|---------|-------------|
| --batch-size <n> | — | Embedding batch size |
Workloads
A workload is a saved job configuration, created with --save <name> on badgr run. Rerun by name instead of retyping all the flags; Badgr tracks success rate, average cost, and the last known-good route.
badgr run . --cmd "python train.py" --max-cost 10 --save my-training-job
badgr workload run my-training-job
badgr workload list
badgr workload info my-training-job
badgr workload delete my-training-job| Subcommand | Description |
|------------|-------------|
| list [n] | List saved workloads |
| info <name> | Stats, route history, recent jobs |
| run <name> | Submit a new job using the workload's saved config (--max-cost, --max-runtime, --set KEY=VALUE to override) |
| delete <name> | Delete the workload record |
Workspaces
Most users never need this — badgr run . handles upload, caching, and artifact storage automatically. Workspaces are for grouping jobs under a named cost/context bucket, optionally linked to S3/GCS storage.
badgr workspace create my-project --storage s3://my-bucket/runs --desc "nightly evals"
badgr run . --cmd "python eval.py" --workspace my-project --max-cost 5
badgr workspace info my-project
badgr workspace list
badgr workspace delete my-projectbadgr batch — generic containerized batch jobs
badgr batch run workload.yml
badgr batch status dep-abc123
badgr batch artifacts dep-abc123
badgr batch receipt dep-abc123
badgr batch compare dep-abc123 dep-def456
badgr batch compare dep-abc123 dep-def456 --key accuracy --higher-is-betterbatch compare reads each run's success_metric by default; --key <metric> compares a different field from the receipt instead, and --higher-is-better (default) / --higher-is-better false controls which run is reported as the winner.
For CV/video/scientific batch, simulation, and physical-AI eval workloads — runs a container from a workload.yml spec and captures output artifacts automatically.
Fan-out — run the same program once per file in a directory, one deployment per input, in parallel:
badgr batch run workload.yml --fan-out ./scenarios
badgr batch run workload.yml --fan-out ./scenarios --max-concurrency 10
badgr batch run workload.yml --fan-out ./scenarios --only failed1.json,failed2.jsonworkload.yml must declare exactly one inputs: entry (the path that varies per task).
| Flag | Default | Description |
|------|---------|-------------|
| --fan-out <dir> | — | Run once per file in this directory instead of a single job |
| --max-concurrency <n> | 5 | Cap in-flight fan-out deployments |
| --only <f1,f2> | — | Rerun just the named input files |
| --dry-run | — | Preview the batch/fan-out plan without provisioning |
badgr sbatch — existing Slurm scripts
badgr sbatch job.slurm
badgr sbatch job.slurm --dry-run
badgr sbatch array_job.slurm # #SBATCH --array=1-100 fans out into one deployment per task
badgr sbatch array_job.slurm --max-concurrency 10Translates --cpus-per-task/--mem/--gres/--time/--export from a real .slurm file. #SBATCH --array=... directives fan out into bounded-concurrency deployments the same way batch run --fan-out does; unsupported directives (--partition, --qos, --account) print a visible warning but don't block translation of the rest.
| Flag | Default | Description |
|------|---------|-------------|
| --max-concurrency <n> | 5 | Cap in-flight array tasks |
| --dry-run | — | Preview the translated job without provisioning |
badgr models — GPU catalog and pricing
badgr modelsLists available GPU types cheapest-first, with VRAM and hourly rate — pulled live from your account when logged in, falling back to the local catalog otherwise. No flags.
badgr template — pre-built workload templates
badgr template list
badgr template info axolotl
badgr serve template vllm --model meta-llama/Llama-3.1-8B-Instruct --max-cost 10
badgr run template axolotl --config ./config.yaml --max-cost 10Provider-neutral templates for common frameworks (vllm, invokeai, comfyui, axolotl, unsloth, …). template list/template info <name> just browse the catalog; launching always goes through badgr serve template <name> or badgr run template <name>, which apply the template's default flags before handing off to the normal serve/run path.
badgr capacity — check live availability
badgr capacity # cheapest runnable GPU across all types
badgr capacity --gpu A100 # a specific GPU type
badgr capacity --gpu A100 --region EU # region-filtered
badgr capacity --gpu A100 --max-price 2.50 # price-cappedbadgr billing
badgr billing status # current balance
badgr billing add 20 # add funds — $5 minimum top-upRequirements
- Node.js 20.10+
- A Badgr account — sign up at aibadgr.com
