telos-kit
v2.0.11
Published
Spec-driven capability orchestration for Codex and Claude Code
Readme
Telos V2
A spec-driven harness that keeps implementation and evaluation looping until an approved goal is verified.
Telos is an open-source development harness for Codex and Claude Code. It turns a human goal into a concrete contract and manages implementation, testing, and evaluation until the contract is verified.
Telos does not replace specialist skills or run its own keyword router. The active coding harness selects an implementation capability. Telos keeps that work tied to the frozen Spec, auditable run state, deterministic verification, and an independent final decision.
Why Telos
Generated code alone is not a completed change. A completed change needs a clear goal, bounded scope, tests, evidence, and a reliable decision about whether it is actually done.
Human goal → Spec → Harness-selected capability → Implement → Eval
↑ │
└── Retry ─┘The user owns the goal. Telos owns the workflow, state, and final evaluation.
Harness-selected capabilities
Installed skills, plugins, and agents are optional, replaceable capabilities. Telos does not depend on a specific provider and does not score capability names with keywords.
- The current Codex or Claude Code harness selects a suitable capability from the Spec and repository context.
- Telos records the selected capability as audit information; it does not run a second competing selector.
- Without a suitable external skill, Telos continues with the current Codex or Claude Code session and its built-in tools.
- After a rejected Eval, the harness may select a different capability for the next iteration.
Install
Node.js 20+ is required.
npx telos-kit install all
# or
npm install -g telos-kit
telos install allUse codex or claude instead of all to install one integration. Restart the relevant client after installation.
Remove only Telos-owned plugin files and catalog entries with telos uninstall codex|claude|all.
Update
# npx
npx telos-kit@latest update all
# global installation
npm update -g telos-kit
telos update alltelos update codex updates the local Codex marketplace and re-registers Telos. Restart Codex completely before using the updated skills.
telos update claude copies the updated Claude marketplace, but Claude Code runs marketplace plugins from its own cache. After updating, run the following inside Claude Code, then reload plugins (or restart Claude Code):
/plugin marketplace update telos-kit
/plugin update telos@telos-kit
/reload-pluginsIf the plugin still shows an earlier version, use /plugin to verify that telos@telos-kit is installed and enabled.
Workflow
spec → run → eval$spec//telos:speccreates a frozen Feature SPEC at.telos/specs/<slug>/SPEC.mdwith measurable acceptance criteria.$run//telos:runmanages the implementation → test → evaluation loop.$eval//telos:evalapproves, rejects, or marks the result uncertain using evidence from the frozen Spec.
Start every feature run with an explicit slug. Its local state and Eval reports stay separate from other feature contracts.
telos run start --spec payment-flow --project-root . --capability implementation
telos run record --spec payment-flow --project-root . --status approved --summary "all acceptance criteria pass"
telos run status --spec payment-flow --project-root .run records iterations in .telos/runs/<slug>.json and Eval results in .telos/evals/<slug>/<iteration>.md. It stops for user direction when scope must expand, the Spec conflicts with the repository, no suitable capability is available, or the iteration limit is reached.
telos run start records its slug in the developer-local, Git-ignored .telos/active. The installed hooks use that pointer to gate only paths declared by .telos/project.yml against the active Feature SPEC.
Project verification and scope evidence
.telos/project.yml keeps repository-specific checks out of prompts. Optional risks run only for matching changed paths; scopes declares coverage dimensions used by acceptance criteria.
modules:
- name: web
paths: ["src/**"]
verify: ["npm test"]
risks:
- id: secret-literal
when: ["src/**"]
check: grep -rnE '(token|secret)=' src/
fail_when: found
scopes: ["ios", "android"]For a scoped criterion, provide evidence for every scope, then validate its presence. Telos Eval still decides whether that evidence is sufficient.
- [ ] AC1 [scopes: ios, android] Login succeeds.
- Evidence [ios]: iOS end-to-end test passed.
- Evidence [android]: Android end-to-end test passed.telos verify --changed --project-root .
telos evidence check --spec .telos/specs/payment-flow/SPEC.md --project-root .Projects can make mechanical verification deterministic with a committed .telos/project.yml. telos verify --changed combines tracked changes from HEAD with untracked, non-ignored files, selects every matching module, then runs only that module's declared commands. An empty verify list requires manual verification and cannot pass automatically.
The commands in verify and risks[].check run with your permissions. Review changes to project.yml before executing them, especially changes from an untrusted pull request.
modules:
- name: web
paths: ["src/**"]
verify: ["npm test", "npm run lint"]telos verify --changed --project-root .
telos run status --spec payment-flow --project-root .For first-time setup and diagnostics:
telos init --project-root .
telos doctor --project-root .init creates one paths: ["**"] module and infers at most one verification command from repository markers. It writes verify: [] when detection is uncertain, and adds .telos/runs/, .telos/evals/, .telos/active, and .telos/hook-warnings/ to .gitignore. doctor parses the configuration, reports module path matches, checks every configured command, and warns about Git-tracked Telos runtime state.
History and review
verify --changed stores a worktree fingerprint on an active run. Eval recording refuses a missing or stale snapshot, so code changed after verification cannot be approved accidentally. Reusing a completed slug archives its prior state and Markdown reports.
telos history --since 7d --project-root .$review / /telos:review turns repeated rejection history into report-only prevention proposals. It prefers regression tests, then tool-managed lint/type rules, then an expressible risk check. It never edits project.yml. Optional risk metadata (added, origin, caught) supports this human review; a stale risk with caught: 0 is a review candidate, never an automatic deletion.
Principles
- Spec before execution — work begins with an explicit completion contract.
- Capabilities over providers — choose skills by what they can do, not their name or installation order.
- Evidence over intent — implementation is complete only when evaluation evidence supports the acceptance criteria.
- Safe, bounded loops — do not retry blindly; stop when user judgment or a new capability is required.
- Codex and Claude Code together — the same V2 workflow ships for both environments.
Development
npm ci
npm test
npm run check:release
npm pack --dry-runThe npm package ships the Node CLI and both plugin bundles. GitHub Actions tests Node 20, 22, and 24; GitHub Releases can publish verified versions to npm through Trusted Publishing.
