codex-dev-team
v0.1.0
Published
Install a five-role collaborating AI development team for Codex.
Maintainers
Readme
Codex Dev Team
Turn one software request into a real five-role AI development team inside Codex.
Codex Dev Team does not ask one agent to pretend to be five personas. JARVIS creates four named Codex SubAgents, gives them separate responsibilities, and makes them communicate directly through evidence-backed handoffs.
You act as the CEO. You speak only with JARVIS and receive one final report.
Quick start
Install the skill globally for Codex:
npx --yes codex-dev-team@latestThen start a mission:
Use $codex-dev-team to build this feature.JARVIS inspects the project, asks at most three material questions, and presents Mission Brief M1. The team does not edit anything until you confirm it.
When to use it
Codex Dev Team is intentionally heavier than a normal single-agent workflow. Use it when independent planning, implementation, review, and adversarial testing justify the extra coordination:
- multi-file features and architectural changes;
- risky refactors or migrations;
- bugs with unclear causes or expensive regressions;
- authentication, payments, permissions, or data-flow changes;
- release-critical work that needs an explicit ship recommendation.
For tiny edits, a normal Codex task will usually be faster.
About the demo: the public demo builds a basic Pomodoro timer so the complete agent workflow is easy to understand in a short video. That example is deliberately simple—and deliberately overkill. The team is designed for substantially harder software missions.
The team
| Agent | Role | Responsibility | Project access | |---|---|---|---| | JARVIS | Manager | Confirms the mission, coordinates the team, resolves scope, reports to the CEO | Read-only by default | | DAEDALUS | Architect | Creates the architecture, steps, risks, and acceptance criteria | Read-only | | ATHENA | Reviewer | Challenges the plan and diff, then gives JARVIS an independent ship recommendation | Read-only | | VULCAN | Coder | Implements the approved plan and writes regression tests | Read/write | | NEMESIS | Tester | Tries to break the result and verifies every fix | Read-only |
All five roles run on Codex. No external model is required.
How they collaborate
flowchart TD
CEO["You — CEO"] <--> J["JARVIS — Manager"]
J --> D["DAEDALUS — Architect"]
J --> A["ATHENA — Reviewer"]
J --> V["VULCAN — Coder"]
J --> N["NEMESIS — Tester"]
D <--> A
D <--> N
D <--> V
V <--> A
V <--> N
A -->|"Independent executive review"| J
J -->|"CEO Brief"| CEOThe specialists do not communicate only through JARVIS:
- DAEDALUS sends the plan directly to ATHENA and NEMESIS.
- ATHENA challenges DAEDALUS until the plan is reviewable.
- VULCAN asks DAEDALUS for architectural decisions.
- VULCAN sends completed steps directly to ATHENA and NEMESIS.
- NEMESIS sends reproducible defects directly to VULCAN.
- VULCAN responds with a fix or repeatable counter-evidence.
- NEMESIS independently retests.
- ATHENA sends the final independent review directly to JARVIS.
- JARVIS reports the result to you as CEO.
Visible Codex SubAgents
JARVIS is the main Codex agent. The other four appear as named SubAgent tasks:
daedalus
athena
nemesis
vulcanSome Codex environments limit simultaneous agents. Codex Dev Team therefore runs in visible waves instead of pretending all five are concurrent:
Wave A: JARVIS + DAEDALUS + ATHENA + NEMESIS
Wave B: JARVIS + VULCAN + ATHENA + NEMESIS
Wave C: final specialist closeouts + ATHENA review to JARVISEvery specialist still runs and remains available for reactivation. If native SubAgents are unavailable, the skill stops rather than simulating the team.
Mission lifecycle
Your idea
→ JARVIS asks up to 3 questions
→ Mission Brief
→ Your confirmation as CEO
→ DAEDALUS plan
↔ ATHENA and NEMESIS challenge it
→ VULCAN implementation
↔ ATHENA review
↔ NEMESIS break/fix/retest loop with VULCAN
→ ATHENA executive review to JARVIS
→ JARVIS CEO Brief to youNothing is implemented before you confirm the Mission Brief.
Evidence wins
Every important handoff carries a mission ID, plan version, item ID, claim, evidence, requested action, and blocking status.
A NEMESIS defect must include:
- stable defect ID and severity;
- affected acceptance criterion;
- exact reproduction;
- expected and actual behavior;
- command, output, path, or other evidence.
VULCAN may respond FIXED, REJECTED_WITH_EVIDENCE, or NEEDS_DECISION. NEMESIS then returns VERIFIED, CLOSED_NOT_REPRODUCIBLE, or REOPENED.
The loop is bounded. Unresolved defects remain visible and prevent fake success.
ATHENA and the CEO Brief
ATHENA must send JARVIS one independent final recommendation:
SHIP
SHIP_WITH_RISKS
DO_NOT_SHIPJARVIS cannot hide or soften that verdict. JARVIS turns the team's evidence into a concise CEO Brief containing:
- mission status;
- what was completed;
- ATHENA's recommendation;
- NEMESIS's verification status;
- material risks and disagreements;
- decisions required from you;
- one recommended next move.
Final report
The final result is one native Markdown file that opens directly in Codex:
outputs/<mission>-codex-dev-team-report.mdIt contains the mission, real team activity, plan execution, changed files, acceptance evidence, tests, ATHENA review, NEMESIS defect lifecycle, direct collaboration ledger, open risks, and next step.
Use
Use $codex-dev-team to build this feature.Example:
Use $codex-dev-team to migrate our authentication flow to passkeys without breaking existing sessions.Restart Codex after the first installation if the skill does not appear immediately.
Important limitations
- Multi-agent work uses more time and tokens than a single-agent task.
- Five roles may run in waves because concurrency varies by Codex environment.
- Read-only role boundaries are workflow rules unless the host provides hard sandbox isolation.
- The team preserves user changes and cannot perform destructive or external actions without the usual authorization.
- Final quality still depends on the repository, available tests, tools, and evidence.
Development
npm run demo
npm testLicense
MIT
