llm-armory
v0.1.1
Published
Choose the right executor for the job — named loadouts (grok-high, quality, …) for hybrid Claude advisor + Grok executor fleets.
Downloads
21
Maintainers
Readme
llm-armory
llm-armory — choose the right executor for the job.
The armory holds named loadouts (executor lanes) you can pull for different kinds of work.
Fable (Claude Code) acts as the advisor. When you need heavy lifting, you explicitly arm the session by choosing the right loadout from the armory — primarily grok-high (Grok 4.5 at effort high; Grok 4.5 supports high|medium|low only — no xhigh).
This is built for deliberate hybrid advisor + executor patterns from Claude CLI. Pure Grok sessions should use their native spawn_subagent tools instead.
Usage
armory --list # list available loadouts
armory --dry-run grok-high
armory quality # native Fable/Max advisor session (unarmed)
armory grok-high -p "task brief" -w my-exec # arm with Grok 4.5 @ effort high
armory fleet grok-high --manifest fleet.txt # parallel fleet of children
armory fleet-status # dashboard of fleet worktrees
armory fleet-report # RESULT summary (exit 1 if any bad)The launcher lives at bin/llm.
Install
npm install -g llm-armory # or: npx llm-armory …
# bins: armory | llm-armory | llm
armory --list
armory doctorFrom a git checkout (dev):
mkdir -p ~/.local/bin
ln -sfn "$PWD/bin/llm" ~/.local/bin/armory
export PATH="$HOME/.local/bin:$PATH"
armory --listLoadouts ship with the package (presets/). Override with LLM_ARMORY_HOME if you keep a custom arsenal.
Loadouts in the armory
| Loadout | Backend | Use for |
|-------------|--------------------------|---------|
| quality | Max / Fable (native) | pure advisor / judgment sessions (equivalent to plain claude) |
| grok-high | Grok 4.5 (effort high) | Primary loadout — arm Fable advisor sessions. Pins --model grok-4.5 --effort high + worktrees + contract. Grok 4.5 efforts: high|medium|low only. |
| grok-xhigh | alias → same as grok-high | Deprecated name kept for old prompts; does not unlock a stronger effort. |
| balanced | DeepSeek API | (skipped for now) |
| glm | z.ai API | (skipped for now) |
| free | freellmapi (self-hosted) | (skipped for now) |
| burn | Anthropic API key | limit-reset days, uncapped Opus |
Current focus: Fable as advisor + armory grok-high. Legacy cheap dispatch is disabled; grok-high is the ready Grok 4.5 loadout.
Skills
The armory ships an optional Claude Code skill.
fusion-advisor — two-model advisor fusion
When a decision is high-leverage (a design fork, a risky change, a /loop course-correction), a single frontier model can be confidently wrong. fusion-advisor turns your advisor session (Claude Code / Opus) into a two-model advisor: it consults a second, different-vendor frontier model (Grok, via the armory launcher) as a context-isolated peer, reconciles the two positions by verifying — not voting, then delegates execution down to the armory's executor loadouts.
It is built on the research reality that naive multi-model mixing often hurts (quality dilution, echo-chamber convergence, self-preference bias) — so the skill is mostly guardrails that make fusion pay off only where it should. Manual-invoke only; reserved for high-leverage calls (it costs 2–10× tokens).
- Skill file:
skills/fusion-advisor/SKILL.md - Install:
ln -sfn "$PWD/skills/fusion-advisor" ~/.claude/skills/fusion-advisor
Fleet
When one-off armory <loadout> -w x -p "..." is not enough, fleet launches
many children from a manifest (one worktree + prompt per line), with
--max-parallel, --stagger, optional --seed file copies, and bookkeeping
under each worktree (.child-out.log, .child-pid, .child-exit).
Manifest format (# comments and blank lines skipped):
# name|prompt-file (paths relative to the manifest's directory)
fix-auth|prompts/fix-auth.md
fix-billing|prompts/fix-billing.mdWorked example:
# From the target app repo (or pass --cwd):
cat > /tmp/fleet.txt <<'EOF'
child-a|prompts/a.md
child-b|prompts/b.md
EOF
armory fleet grok-high --manifest /tmp/fleet.txt \
--max-parallel 10 --stagger 3 --seed .env
armory fleet-status --cwd .
# child-a | running(pid 12345) | commits=0 | task 1: done — ...
# child-b | exit=0 | commits=2 | task 3: done — ...
armory fleet-report --cwd .
# child-a | exit=0 | commits=2 | RESULT: ok — commits: 2 — fixed auth
# child-b | exit=0 | commits=1 | RESULT: ok — commits: 1 — billing edge
# fleet: 2 ok · 0 bad of 2Each child lands in <repo>/.claude/worktrees/<name> (same layout as -w).
If that path already exists, fleet refuses that child loudly and continues
others (no silent reuse). Launch success is separate from child success:
fleet exits 0 when every child was launched; use fleet-report as the gate.
Executor contract
Every loadout pulled from the armory carries strong discipline:
- One commit per completed task (no batching)
PROGRESS.mdledger written after each task (left untracked)- Session ends with exactly one line:
RESULT: ok|partial|failed — commits: <n> — <summary>
These rules are injected via --rules (Grok) or system prompt (other lanes).
Conventions
- The advisor session (native Max/Fable) never sets
ANTHROPIC_BASE_URL. - Cross-CLI delegation: Fable (advisor) explicitly spawns
armory <loadout> -p "..."children. - Keys live in
presets/providers/*.env(gitignored).
Statusline
Show which loadout is currently armed:
"statusLine": {
"type": "command",
"command": "~/Repositories/llm-armory/bin/llm-statusline"
}Notes
- Designed primarily for Claude Code CLI (Fable as advisor) to arm itself with Grok 4.5 (
grok-high) executors on explicit instruction. - Grok-native sessions are unaffected — use
spawn_subagentand native tools. - The legacy "cheap LLM pool" dispatch remains disabled.
See templates/armory-snippet.md for a ready-to-paste block for your CLAUDE.md / AGENTS.md.
Works well with
The armory is standalone-first: loadouts launch the same way with or without
sibling tools installed. Nothing errors when they are absent. When co-installed,
optional extras appear via the Agent Status Providers
convention (armory writes launch records; details in INTEROP.md).
| Sibling | Extra when installed |
|---------|----------------------|
| status-herald | Curtain / bars can show model@effort (and preset) for armory children from the launch session stamp. |
| agentic-sage | Nested fleet provenance via SAGE_PARENT (exported at launch) and the session record’s parent_session field. |
| token-oracle | Refreshes live model truth after launch; armory only stamps what was known at exec. |
Credits
The free loadout runs against freellmapi — a self-hosted free-tier LLM API aggregator by tashfeenahmed (this project is not affiliated with it). Self-host runbook: freellmapi/NOTES.md.
Community
- Website: armory.muslewski.com
- Questions & ideas: Discussions
- Bugs & features: Issues
- Contributing: CONTRIBUTING.md
- Code of Conduct: CODE_OF_CONDUCT.md
- Security: SECURITY.md (private reports only)
- Support matrix: SUPPORT.md
If you're not sure whether something is a bug, start a Discussion — maintainers can promote it to an issue when it is.
