@rrrtx/xr
v1.0.0
Published
XR — a local-first, provider-neutral AI agent runtime. BYOK, spend-capped, tamper-evident audit, plugin/MCP extensibility.
Maintainers
Readme
XR
An AI agent runtime you can actually audit.
Give it a task. It plans, uses tools, and changes real things on your machine — under a policy gate, your approval, a spend ceiling, and a hash-chained log you can verify offline.
Quick start · What XR is · How it works · Providers · Security · Docs · Contributing
Version: 1.0.0 (Truth) · Package: @rrrtx/xr · License: MIT
Status: Public Beta. @rrrtx/xr is honestly labeled beta software: install and use it, expect the documented golden path to work on the validated platforms, and check the support matrix and known-limitations register before adopting it for anything critical.
v*-beta.*tags land on the prerelease channel (npmbetadist-tag, GitHub prerelease) for early adopters; feedback goes through the beta loop.
Version source of truth:
release.manifest.json. Every surface —src/core/version.ts,package.json, this README,install.sh,install.ps1and the website — is stamped from that one file, and CI fails the build if any of them drift (Constitution Article XXII.1).
XR — a local-first, provider-neutral AI agent runtime. BYOK, spend-capped, tamper-evident audit, plugin/MCP extensibility.
Bundled skills: 65 (counted from skills/ at release time.)
What is XR, in plain language?
You type a task in your terminal. XR figures out the steps, calls a language model, and uses tools — reading and writing files, running commands, browsing, calling APIs — until the task is done or it honestly reports that it failed.
The difference is what surrounds that loop:
| | | |---|---| | 🔑 You bring the key | XR ships no API key and no cloud account. It runs on your provider key or a model running on your machine. | | 🛑 It asks before it acts | Consequential actions stop and wait for your approval. Denial is enforced in the execution path, not suggested in a prompt. | | 💸 It cannot overspend | Every task carries a USD and token ceiling, checked during the loop, not after the bill. | | 🔗 It writes down what it did | Every event is SHA-256-linked into a local chain you can verify offline with one command. | | 💻 It runs on your machine | One SQLite database, no telemetry, no mandatory network. Ten local model runtimes are first-class. |
xr "summarize the open TODOs in this repo and draft a cleanup plan"Who it's for: developers automating real work on real machines; teams that need an agent's actions to be reviewable after the fact; anyone who wants an agent that runs fully offline.
What XR is — and what it is not
XR is:
- a self-hosted agent runtime — no mandatory cloud, no telemetry;
- provider-neutral — 26 presets, 16 hosted (BYOK) + 10 local runtimes, one switch command;
- governed — policy, approvals, budgets and audit are enforced in the execution path, not promised in docs;
- extensible — skills, plugins and MCP servers reachable identically from every surface;
- MIT-licensed and readable end to end.
XR is not:
- not certified against SOC 2, ISO 27001, HIPAA, PCI-DSS or FedRAMP — no external audit exists;
- not a sandbox — in-process policy enforcement, not kernel or VM isolation;
- not a hosted product — there is no XR cloud;
- not a substitute for a human reviewing consequential actions;
- not finished — the known-limitations register is a first-class release artifact, and the release docs state precisely what beta means today (support matrix).
Every capability claim on this page is backed by evidence recorded in
release.manifest.jsonand re-checked in CI bybun run claim-lint, which also fails the build on a list of prohibited overclaims. If a sentence here cannot be evidenced, CI rejects it.
Quick start
1. Install
| Channel | Platform | Command |
|---|---|---|
| Binary (default) | Linux · macOS · Termux · WSL | curl -fsSL https://raw.githubusercontent.com/ahmadrrrtx/xr/main/install.sh \| bash |
| Binary (default) | Windows PowerShell 5.1 / 7+ | iex (irm https://raw.githubusercontent.com/ahmadrrrtx/xr/main/install.ps1) |
| Homebrew | macOS · Linux | brew install ahmadrrrtx/tap/xr |
| WinGet | Windows | winget install ahmadrrrtx.XR |
| Scoop | Windows | download scoop/xr.json from the release · scoop install ./xr.json |
| .deb | Debian · Ubuntu | download xr_<ver>_amd64.deb · sudo dpkg -i xr_*_amd64.deb |
| Docker | any | docker run ghcr.io/ahmadrrrtx/xr:latest |
| From source | any | git clone https://github.com/ahmadrrrtx/xr && cd xr && bun install |
⚠ npm
latestis still the 3.x line. The badges above are dual on purpose:latest (legacy)is3.1.5(pre-rebaseline) andbetais empty until Phase 3 publishes1.0.0-beta.1.bun add -g @rrrtx/xrtherefore installs the old build. Because3.1.5sorts higher than1.0.0, the first stable 1.0.0 publish must re-pointlatestexplicitly (npm dist-tag add @rrrtx/[email protected] latest) — a beta must not move that tag. See the release runbook, the version ladder, and known limitations. Until then use the binary channel or build from source.
Every channel installs the same canonical build. Tagged releases ship cosign keyless signatures
over SHA256SUMS, a CycloneDX SBOM and SLSA3 provenance — verify with
docs/release/VERIFYING_RELEASES.md. Channel configs are
generated from the release manifest and drift-gated (bun run channel:check), so a channel
cannot fall behind the release it serves. Publication status per channel:
docs/release/SUPPORT_MATRIX.md.
2. First run
xr onboarding # guided setup (provider + memory + optional voice)
xr doctor # health check — exits non-zero if XR cannot actually workxr doctor answers exactly one question: can XR run a task right now? It exits non-zero
when no provider is reachable and tells you the single next action. It never prints ok for a
system that cannot do work.
3. Your first task
xr "hello, XR" # one-shot task
xr # fullscreen interactive shell
xr serve # dashboard + chat at http://localhost:3141 (127.0.0.1, token-authed)New here? Follow the golden path end to end:
docs/development/GETTING_STARTED.md.
Run it fully offline
ollama serve && ollama pull qwen2.5:7b # any of 10 supported local runtimes
xr providers set ollama qwen2.5:7b
xr "refactor this function" # no network requiredHow XR works
Every surface — the CLI, the shell, Telegram, the daemon's chat — funnels into one call. There is no side door: policy, approvals, budget, cancellation and audit are properties of the pipeline, so an interface cannot skip them.
flowchart TB
subgraph S["SURFACES"]
direction LR
CLI["xr task<br/><small>src/commands</small>"]
SH["Shell + TUI<br/><small>src/interfaces</small>"]
TG["Telegram<br/><small>src/telegram</small>"]
DA["xr serve<br/><small>src/daemon · 127.0.0.1</small>"]
end
S -->|"AgentService.execute(request)"| EX
EX["<b>EXECUTION FABRIC</b> · src/execution<br/>envelope · state machine · idempotency keys<br/>leases · checkpoints · runner = sole loop caller"]
EX --> LOOP["<b>AGENT LOOP</b> · src/core/agent<br/>chat → tools → observe · plan/act hybrid<br/>turn-repair (strict JSON) · memory writes"]
LOOP --> PR["<b>PROVIDERS</b><br/>src/providers<br/>26 presets · 5 native<br/>adapters + OpenAI-compat<br/>health() + failover"]
LOOP --> TL["<b>TOOLS</b><br/>src/tools<br/>files · git · shell<br/>browse · guarded registry"]
LOOP --> ME["<b>MEMORY & CONTEXT</b><br/>src/context<br/>retrieval · embeddings<br/>compression · summaries"]
PR --> TP
TL --> TP
ME --> TP
TP["<b>TRUST PLANE</b> — cross-cutting, same pipeline<br/>policy gate · approvals · budget governor · egress allowlist<br/>secrets vault · hash-chained audit log"]
TP --> ST["<b>LOCAL STATE</b> · src/state<br/>SQLite workspace store · migrations · write-gate · repos"]
style EX fill:#0b1220,stroke:#00d2ff,color:#e6f6ff
style LOOP fill:#0b1220,stroke:#00d2ff,color:#e6f6ff
style TP fill:#1a0f2e,stroke:#9a6bff,color:#f0e6ff
style ST fill:#0b1220,stroke:#4a5568,color:#e6f6ffWhat happens to one task
flowchart TD
R["request"] --> E["AgentService.execute<br/><small>envelope · state machine · idempotency · audit seed</small>"]
E --> RUN["RUNNER<br/><small>the only code allowed to drive the loop</small>"]
RUN --> C0{"⓪ cancelled?"}
C0 -->|yes| CAN["outcome: cancelled<br/><small>honest, not a fake completion</small>"]
C0 -->|no| C1["① chat completion<br/><small>provider failover · tokens metered</small>"]
C1 --> C2{"② turn contract<br/>strict JSON valid?"}
C2 -->|no| REP["turn-repair · src/reliability<br/><small>else the step fails</small>"]
REP --> C2
C2 -->|yes| C3{"③ cancelled?"}
C3 -->|yes| CAN
C3 -->|no| T["④ for each tool call"]
T --> A{"approval required?"}
A -->|denied| TE["tool error — never executed"]
A -->|granted / not needed| P{"policy gate<br/><small>risk class · egress allowlist</small>"}
P -->|blocked| TE
P -->|allowed| B{"budget governor<br/><small>USD + tokens</small>"}
B -->|exceeded| STOP["outcome: budget stop"]
B -->|within cap| EXEC["execute tool → observation"]
EXEC --> M["⑤ memory delta + provenance"]
TE --> M
M --> D{"done?"}
D -->|no| C0
D -->|yes| OK["outcome: success | failed"]
OK --> AUD["audit chain: every event SHA-256-linked<br/><small>verify offline: xr audit verify</small>"]
STOP --> AUD
CAN --> AUD
style CAN fill:#2e1a1a,stroke:#ff6b6b,color:#ffe6e6
style STOP fill:#2e2a1a,stroke:#ffd93d,color:#fff9e6
style OK fill:#0f2e1a,stroke:#51cf66,color:#e6ffe6
style AUD fill:#1a0f2e,stroke:#9a6bff,color:#f0e6ffA run ends success, failed, or cancelled — never a fake completion. Cancellation is
cooperative and real: in the Shell, Ctrl+C/Esc stop the current run (a pending approval is
denied fail-closed first); xr run maps the first SIGINT to a cooperative wrap and exits
130, a second forces exit. When an action genuinely cannot be interrupted, XR stamps the
honest outcome instead of claiming it stopped cleanly.
Design principles
- One computation authority per question. Whatever answers a question for you answers it
for every surface and for CI — doctor's readiness engine is the onboarding capability scan;
the cross-platform suite is one file list (
scripts/platform-parity.ts) executed per OS; channel configs are generated, never handwritten. - No bypass around the runner. Every turn flows through one envelope → runner → loop pipeline, so governance cannot be skipped by an interface.
- Honest outcomes. Never a fake completion.
- State you can inspect. One SQLite database, hash-chained audit events, exportable sessions, reproducible inventory.
- Fast path stays fast. Commands boot only the subsystems they need, the hot path performs zero synchronous FS/process I/O (lint-enforced), and budgets are CI-gated.
Providers
XR ships 26 built-in provider presets — 16 hosted and 10 local runtimes. Swap anytime, no restart, no re-config.
flowchart LR
A["agent loop"] --> REG["provider registry<br/><small>src/providers/registry.ts</small>"]
REG --> NAT["native adapters<br/><small>Anthropic · Google · Mistral<br/>Cohere · AWS Bedrock</small>"]
REG --> OAI["OpenAI-compatible transport<br/><small>src/providers/openai-compat.ts</small>"]
OAI --> H["11 hosted presets<br/><small>OpenAI · Groq · DeepSeek · Cerebras<br/>Together · Fireworks · SambaNova<br/>HuggingFace · OpenRouter · xAI · Perplexity</small>"]
OAI --> L["10 local runtimes<br/><small>Ollama · LM Studio · llama.cpp · Jan<br/>LocalAI · vLLM · GPT4All · KoboldCPP<br/>Text-Gen-WebUI · SGLang</small>"]
OAI --> CU["any OpenAI-compatible base URL"]
REG -.->|"health() probe<br/>+ failover"| A
style L fill:#0f2e1a,stroke:#51cf66,color:#e6ffe6
style NAT fill:#0b1220,stroke:#00d2ff,color:#e6f6ffFive hosted providers use dedicated native API adapters (Anthropic, Google, Mistral, Cohere, AWS Bedrock). Every other preset — remaining hosted providers and all ten local runtimes — speaks the OpenAI-compatible protocol through one transport, and a custom preset can point at any OpenAI-compatible base URL.
xr providers list
xr providers set openai gpt-4o-mini
xr providers add claude # key entered masked → OS keychain, else AES-256-GCM sealed file
xr providers test # probe configured providers liveSwitching models runs a preflight → canary → swap → verify state machine that rolls back
automatically if the new model cannot be reached (--force skips the probe).
Provider count is not a measure of product quality and is deliberately not scored by
xr evaluate. Counted fromPRESETSinsrc/providers/presets.ts.
| Provider | Type | Default model |
|---|---|---|
| Ollama | Local | qwen2.5:7b — auto-detect, model pull, free |
| LM Studio, llama.cpp, Jan, LocalAI, vLLM, GPT4All, KoboldCPP, Text-Generation-WebUI, SGLang | Local | picked per install |
| Claude (Anthropic) | Hosted · native adapter | claude-3-5-sonnet-20241022 |
| Gemini (Google) | Hosted · native adapter | gemini-1.5-flash |
| Mistral | Hosted · native adapter | mistral-small-latest |
| Cohere | Hosted · native adapter | command-r-plus-08-2024 |
| AWS Bedrock | Hosted · native adapter | claude-3-sonnet |
| OpenAI | Hosted | gpt-4o-mini |
| Groq | Hosted | llama-3.3-70b-versatile |
| DeepSeek | Hosted | deepseek-chat |
| Cerebras | Hosted | cerebras/csm-8b |
| Together AI | Hosted | Llama-3.3-70B-Instruct-Turbo |
| Fireworks | Hosted | llama-v3p1-70b-instruct |
| SambaNova | Hosted | Llama-3.1-70B-Instruct |
| Hugging Face | Hosted | Llama-3.1-8B-Instruct |
| OpenRouter | Hosted | anthropic/claude-3.5-sonnet |
| xAI | Hosted | grok-2-latest |
| Perplexity | Hosted | llama-3.1-sonar-large-128k-online |
| + any OpenAI-compatible endpoint | Local/Hosted | your base URL |
Default models are the presets' launch defaults (defaultModel), not ceilings.
Security — the trust plane
flowchart TB
M["model output<br/><small>untrusted data, never instructions</small>"] --> G
subgraph G["TRUST PLANE — enforced inside the execution path"]
direction TB
AP["approvals · src/control/approvals.ts<br/><small>human consent, per workspace, auditable</small>"]
PO["policy gate · src/security/guard.ts<br/><small>risk classes over every tool effect</small>"]
BU["budget governor · src/cost/governor.ts<br/><small>USD + token ceilings, checked mid-loop</small>"]
EG["egress allowlist · src/security/egress-proxy.ts<br/><small>only configured domains receive data</small>"]
AP --> PO --> BU --> EG
end
G -->|allowed| EFF["real effects<br/><small>files · shell · network · desktop</small>"]
G -->|"blocked / denied / over budget"| REJ["refused + recorded"]
EFF --> AUD["hash-chained audit log<br/><small>src/state/workspace-store.ts</small>"]
REJ --> AUD
AUD --> V["xr audit verify<br/><small>offline verification</small>"]
style G fill:#1a0f2e,stroke:#9a6bff,color:#f0e6ff
style REJ fill:#2e1a1a,stroke:#ff6b6b,color:#ffe6e6
style V fill:#0f2e1a,stroke:#51cf66,color:#e6ffe6XR Shield — the enforcement boundary
XR Shield is the name for the seven components every consequential action passes
through. It is not a product tier and not a scanner: it is the code that can say no.
Enumerable in one import — src/xr-shield/index.ts — so it can be audited as a unit.
| # | Component | Where | Question it answers | Can refuse? |
|---|---|---|---|---|
| 1 | Capability policy | src/capabilities/policy.ts | Is this capability permitted for this workspace and mode? | ✅ |
| 2 | Action guard | src/security/guard.ts | Is the action dangerous once fully decoded and canonicalized? | ✅ |
| 3 | Trust lattice + placement | src/runtime/trust/ | What risk tier is this, and where may it run? | ✅ |
| 4 | Consent / approvals | src/control/approval-store.ts | Has a human approved this? (durable, TTL default-deny) | ✅ |
| 5 | Network egress | src/security/egress-proxy.ts, private-ip.ts | Is this destination allowed? | ✅ |
| 6 | Execution + output integrity | src/security/exec-integrity.ts, tool-output.ts, secret-broker.ts | Is this binary allowlisted? Is tool output framed as data before re-entering the prompt? Are secrets brokered, not handed over? | ✅ |
| 7 | Signed audit evidence | src/security/audit-signer.ts, audit-verify.ts | Can this be proven afterwards, and would tampering show? | ➖ records |
This table is generated from XR_SHIELD_COMPONENTS, and
test/architecture/xr-shield-facade.test.ts fails if a listed module stops existing —
the docs cannot drift from the code.
Alongside the boundary, two mechanisms constrain cost and supply chain:
| Mechanism | Where | What it enforces |
|---|---|---|
| Budget governor | src/cost/governor.ts | USD + token ceilings per task, checked before and during steps |
| Secrets at rest | src/config/config.ts | OS keychain where available; else AES-256-GCM sealed file (auto-migrates plaintext); redacted from all status output |
| Plugin trust | src/plugins/ | Manifest permissions, tree hashes, static scan, health checks, disable/remove |
| Prompt injection | src/core/agent.ts | Tool output is treated as untrusted data, never as instructions |
Naming note (Phase 5):
xr shieldused to be a host scanner — processes, startup items, privacy settings. It never governed agent actions, so it is nowxr hygiene, and "XR Shield" names the boundary above.xr shieldstill works with a deprecation notice until 2.0.0. See ADR-0027.
Honesty box: XR enforces in-process policy, not kernel/VM isolation — a guard rail, not a confinement boundary, and not a substitute for reviewing consequential actions. The gaps are written down, not hidden:
docs/security/KNOWN_LIMITATIONS.md.
Reporting a vulnerability: SECURITY.md.
Capabilities
flowchart LR
L["agent loop"] --> RG["tool registry<br/><small>src/tools</small>"]
RG --> SK["<b>Skills</b> · skills/<br/><small>65 bundled · xr-skill.json manifests<br/>loader → registry → resolver → validator</small>"]
RG --> PL["<b>Plugins</b> · src/plugins<br/><small>explicit permissions · hash verification<br/>discover → install → enable → update<br/>→ rollback → quarantine → uninstall</small>"]
RG --> MC["<b>MCP</b> · src/mcp<br/><small>client + manager + registry<br/>stdio · SSE · HTTP · allowlisted</small>"]
RG --> WT["<b>Templates</b> · src/templates<br/><small>11 built-in workflows</small>"]
SK & PL & MC & WT -.->|"same trust plane<br/>approvals · policy · budget · audit"| TP["governed execution"]
style TP fill:#1a0f2e,stroke:#9a6bff,color:#f0e6ffAll four are reachable identically from every surface, and all four are governed by the same trust plane — an MCP server gets no more privilege than a bundled tool.
- Memory (
src/context/): consent-first capture (only what you ask it to remember), categorized + scoped entries, TTL/expiry viaxr memory prune, explainable retrieval (xr memory recall "…"shows match % and why), optional session summaries. - Research (
src/research/): offline by default;xr research deep --allow-public-webopts into live fetching with egress rules and per-run budgets; results carry source provenance. - Voice (
src/voice/,src/interfaces/,xr voice …): optional, local-first adapters; voice can trigger capabilities through the same governed pipeline — approvals still apply. - Computer control (
src/control/,src/computer/): guarded desktop actions behind the approval plane; the vision agent is opt-in; platform support and limits documented per OS.
src/services/multi-agent-service.ts + src/agents/ compile a deterministic plan per
goal kind, then execute it as bounded-concurrency lanes — planner → parallel
worker lanes → reviewer → synthesizer, with a deterministic security gate between
stages and a read-only artifact verifier after synthesis:
xr agents list
xr agents plan "refactor this repo safely"
xr agents run "compare the three vendor pricing plans"
xr run "start a long task" # journaled; crash-safe
xr run --resume <sessionId> # continue from the last durable checkpoint- One funded tree, not N budgets. A workflow gets a single root envelope
(
budget.perTaskUsd/perTaskTokens) that is partitioned across the workers; each step admits against its child slice and the shared root in one ledger transaction, so the tree can never spend more than the root plus one in-flight step. A worker request that tries to carry its own budget is ignored on this path — deliberately: that was the per-worker multiplier this replaced. - Verifiers earn completion. The artifact verifier re-checks claimed files against the workspace (hashes, missing files); its verdict must be strict-JSON with a reason — prose assurance, garbage, or a reject all fail the task.
- Delegation is depth-capped at 1. Workers carry minted identities and can never spawn their own sub-workers; attempted grandchild spawns are refused and audited, not silently flattened.
- Crashes are recoverable, honestly. Per-step checkpoints carry a hash chain and the consumed meter; a tampered chain or a missing checkpoint refuses to resume rather than pretending. A resumed run re-asks the model from the checkpointed transcript — the seed is durable, the model's answer is not.
- The review gate consumes a strict-JSON decision from the deterministic
security_checker; prose-only reviewers fail closed — an unparsable verdict blocks the run rather than waving it through. - Honest failure mapping: transport errors, budget stops and approval stops mark the task failed instead of faking completion.
- Cancellation is workload-aware: stopping a workflow aborts the in-flight worker via a live run map; remaining work is marked, never silently dropped.
Concurrency, per-workflow worker caps, plan-fragment supervision, and verifier
defaults live under orchestration.* in the config. The orchestration plane is
single-process: budgets are enforced against the shared workspace store, not
across machines.
An optional, default-off, effect-verified extension over local-first records. It ships no hosted service, no SLA and no paid tier.
Since Phase 5 it lives outside core as @rrrtx/business-os (ADR-0028) —
core carried 33,759 LOC of enterprise and business surface that no user had installed, and one
maintainer cannot own review, audit and release weight for capability nobody asked for. Behaviour
is unchanged: still default-off, still fail-closed, /api/v1/business/* still answers honestly
when the extension is absent.
bun add -g @rrrtx/business-os # business operating layer
bun add -g @rrrtx/xr-enterprise # org policy, audit export, SLOs, DR, evaluation harnessSee docs/business-os-extension.md and the
migration guide.
Scripting & exit codes
xr doctor --json is the stable machine-readable entrypoint (redacted config, provider
readiness, summary.runnable verdict; secrets are never printed — only presence).
| Code | Meaning |
|---|---|
| 0 | ok |
| 1 | runtime/task failure |
| 2 | usage error |
| 3 | network |
| 4 | denied |
| 5 | not found |
| 130 | interrupted (Ctrl+C / SIGINT) |
Full contract: docs/guides/cli-compat.md.
Performance — budgets, not boasts
Every performance claim is a published budget with a measured baseline and a CI regression
gate (docs/perf/PERF-BUDGETS.md; baseline docs/perf/baseline-1.0.0-source.json):
| Surface | Budget | Measured (1.0.0 baseline) |
|---|---|---|
| --version / --help warm p95 | < 150 ms | 39.8 / 40.5 ms |
| --version / --help cold p95 | < 300 ms | 40.8 / 42.5 ms |
| doctor | < 1 s measured (gate ceiling 2500 ms on shared runners) | 586 ms p95 |
| route decision | < 20 ms | sub-ms |
| dashboard first render | < 1 s | 12.1 ms |
| retrieval @100k items | gate ceiling 250 ms | 24–29 ms |
The fast path performs zero synchronous FS/process I/O (lint-enforced), and a command boots only the subsystems it needs (boot profiles).
How XR proves itself
XR's differentiator is that its claims are checked by machines on every PR.
| Gate | What it pins |
|---|---|
| bun test + parity suite | 2,946 tests across 297 files (707 enterprise/business tests moved to the satellites with their code, ADR-0028); one computation authority (scripts/platform-parity.ts) executed per OS on Linux/macOS/Windows via segmented runs with crash-class retry and file-level culprit attribution |
| release:check + claim-lint | version identity stamped everywhere; every public claim has evidence; prohibited/supervised terms fail the build |
| baseline:inventory | source-derived repository inventory regenerated and compared |
| boundaries + ownership:check + size-gate | layering, area ownership, per-file size discipline (waivers explicit) and a 136,000-LOC ceiling on the whole tree |
| claim-lint constitution gate | every Article N cited anywhere in src/, test/, scripts/, .github/ exists in docs/CONSTITUTION.md |
| api:schema:check + client:check + api:compat | daemon OpenAPI schema, generated client, compatibility |
| channel:check | channel configs match the release manifest |
| supply chain | osv-scanner + bun audit, gitleaks, license scan, SBOM drift, --ignore-scripts hygiene, container scan |
| Quality Gate | single required aggregation check over all of the above |
The evaluation harness moved to @rrrtx/xr-enterprise in Phase 5 (ADR-0028):
bun add -g @rrrtx/xr-enterprise
xr-enterprise evaluate run --offline # 14 suites, 38 scenarios, no network required
xr-enterprise evaluate claims # every public claim mapped to its evidence
xr-enterprise evaluate limitations # what the benchmarks do NOT measure
xr-enterprise evaluate compare <a> <b> # regression detection between releases
xr-enterprise evaluate export <runId> # hash-verifiable evidence bundleIt is outcome-based: a scenario passes only when reality is inspected — an artifact on disk, a
durable record, a state transition, an audit-chain entry. Typing xr evaluate on core prints
where it went and exits 2; the gates in the table above still run in core CI on every PR.
Repository structure
xr/
├─ bin/xr compiled-binary-first launcher (falls back to source)
├─ src/
│ ├─ cli/ router, catalog, lazy command loaders, flags, exit codes
│ ├─ commands/ one file per CLI command (run, doctor, agents, mcp, audit…)
│ ├─ interfaces/ shell + TUI, provider/model pickers, onboarding
│ ├─ core/ kernel: DI container/lifecycle (app.ts), agent loop (agent.ts)
│ ├─ services/ agent-, multi-agent-, budget-, config-, mcp-service
│ ├─ execution/ envelope, runner, state machine, adapters, leases
│ ├─ agents/ multi-agent planner/registry/types
│ ├─ providers/ presets, factory, health, native adapters, openai-compat
│ ├─ tools/ tool registry + guarded tools (files, git, control, egress)
│ ├─ reliability/ turn repair, grammar, profiles
│ ├─ security/ guard, egress-proxy, exec/output integrity, audit signing
│ ├─ xr-shield/ the enforcement boundary, named and enumerable (ADR-0027)
│ ├─ hygiene/ host-hygiene scanner (`xr hygiene`; was `shield`)
│ ├─ state/ SQLite workspace store, repos, write-gate, migrations
│ ├─ context/ memory: assembler, retrieval, embeddings, compression
│ ├─ research/ research engine (offline default, opt-in web)
│ ├─ mcp/ plugins/ skills/ local/ cost/ control/ computer/ telegram/ voice/
│ ├─ daemon/ `xr serve`: API routes, dashboard, chat (127.0.0.1)
│ ├─ update/ install/ atomic updater + channels; install/uninstall
│ ├─ observability/ ui/ metrics/logs/exporters; design system
│ └─ index.ts CLI entry
├─ satellites/ extracted packages, released separately (ADR-0028)
│ ├─ xr-enterprise/ @rrrtx/xr-enterprise — org policy, audit export, SLOs, DR
│ └─ business-os/ @rrrtx/business-os — default-off business operating layer
├─ skills/ 65 bundled skill manifests
├─ plugins/ bundled plugins
├─ scripts/ gates, release machinery, parity runner, perf budgets
├─ test/ 297-file suite mirroring src/ + helpers + fixtures
├─ packaging/ homebrew · winget · scoop manifests (generated)
├─ docs/ product, development, release, security, historical
└─ website/ docs/marketing site (Next.js; scanned by claim-lint)Layering is enforced in CI: the boundaries gate + test/architecture/* pin allowed dependency
directions (surfaces → services → execution/core → state; tools/providers as leaves), and
bun run ownership:check requires every source area to have an owning document. Core imports
nothing from satellites/ — enforced three ways (dependency-cruiser rule, boundary test,
isolation test) — and bun run size-gate holds the whole tree under a 136,000-LOC ceiling so the
23,376 LOC Phase 5 removed cannot quietly grow back.
Uninstall & data
xr uninstall removes the binary, PATH entry and data directory, with explicit flags for what to
keep; the exact matrix is in docs/development/GETTING_STARTED.md.
Updates follow the channel you installed from (brew upgrade xr,
winget upgrade ahmadrrrtx.XR, apt-get install --only-upgrade xr, or xr update for
binary/npm/git layouts) and are atomic with an automatic rollback path.
Documentation map
| Doc | Purpose |
|---|---|
| docs/CONSTITUTION.md | The law the codebase cites — 30 Articles + the Commandments. Enforced: claim-lint fails if code cites an Article that isn't there |
| docs/development/GETTING_STARTED.md | The golden path: install → onboarding → provider → first task → restart/resume → uninstall |
| docs/guides/cli-compat.md | Exit codes, global flags, --yes semantics, scripting envs |
| docs/security/KNOWN_LIMITATIONS.md | Canonical known-limitations register (living) |
| docs/release/SUPPORT_MATRIX.md | Platform/channel support truth per release |
| docs/HISTORY.md | Version ladder (0.2 / 3.x / 4.x / 7.x / 1.x) + the structural ladder of what moved, phase by phase |
| docs/migration/PHASE-5-SATELLITES.md | Phase 5: xr shield→xr hygiene, satellites, the re-based deprecation timeline |
| docs/release/RELEASING.md | The release runbook |
| docs/release/VERIFYING_RELEASES.md | cosign/SBOM/SLSA verification walkthrough |
| docs/release/BETA.md | Beta loop and feedback channel |
| docs/ | Full documentation index |
Contributing
Contributions are welcome — read CONTRIBUTING.md and the
security policy first. The quality bar is identical for humans and agents: every
change ships with tests, passes the local gate battery, and keeps claims honest.
git clone https://github.com/ahmadrrrtx/xr && cd xr
bun install --frozen-lockfile
bun run ci # typecheck + tests + release:check + claim-lint + inventory + gatesAlso: CODE_OF_CONDUCT.md · issue templates ·
docs/development/
Found an inaccurate claim in this README or the docs? That is a bug with its own issue template (false claim) — please file it.
License
MIT — see LICENSE. XR is free software: no paid tier, no telemetry, no lock-in.
