fullstack-agentic-flow
v2.0.3
Published
Gated multi-agent delivery pipeline for Claude Code and Codex: MVC, unified API, and split backend/frontend topologies, TDD-first, typed data, impact analysis.
Maintainers
Readme
fullstack-agentic-flow
A gated, multi-agent delivery pipeline for Claude Code and Codex. Backend and frontend are designed and built in parallel against a frozen interface seam, test-first, with typed data at every boundary, and a measured risk class that decides how much process each change gets.
npx fullstack-agentic-flow initIt works on four repository shapes:
| Topology | What the repo holds | The seam between tracks |
|----------|---------------------|-------------------------|
| mvc | Server-rendered app (Laravel + Inertia/Blade, Rails, Django templates…) | A typed view model per screen |
| unified-api | Backend API + frontend in one repo (REST, GraphQL, RPC) | The API contract, in-repo |
| split-backend | API only; frontends live elsewhere | A contract exported to contracts/ |
| split-frontend | UI only, any framework (React, Vue, Svelte, mobile…) | A contract imported from the API repo |
The agent prompts name no language or framework. Everything repo-specific lives
in one generated file, architecture-context.md. Everything that should be the
same across all your repos (module layout, naming, typed data, TDD) lives in
the pipeline's canon.
What v2 adds
- Topologies. The seam agent (02d) freezes whatever crosses between tracks:
page contracts in MVC, endpoints or GraphQL operations in API repos, and a
versioned contract file that travels between split repos (
/seam export,/seam import). - Backend canon. One module shape for every framework: Domain /
Application / Infrastructure / Interface, one use case per action, fixed
repository method names. There are mappings for Laravel, NestJS, Express,
Django/FastAPI, Spring and Go.
/scaffold-modulegenerates a new module from the repo's mapping, so module eighteen looks like module two. - Typed data (T1–T6). No untyped arrays, maps,
mixedoranyacross a boundary. Named Data classes instead, even when that means many of them. Contracts name every shape, and a diff-based CI check enforces T1 for PHP, TypeScript and Python. - TDD. Contracts name the tests, with the reason each should fail. Sequences put the acceptance tests first as pending specs. Implementers run red → green → refactor per test and log it. CI checks that source changes come with test changes.
- Impact analysis (01b). Before design, the agent predicts the blast radius with the code graphs and scores ten dimensions. The highest score sets the risk class L0–L3, and L2/L3 add characterisation tests, a mandatory perf review, and rollout plans. After implementation it compares actual against predicted and flags untested escapes before merge.
- Toolchain. find-skills, superpowers, claude-mem, impeccable,
task-observer, ponytail, headroom, code-review-graph and graphify, each wired
to specific agents under an explicit precedence order
(
.ai-agents/toolchain.md). - Nothing after the merge. QA, security and performance review run against
the branch diff, so Critical findings block the merge rather than landing on
main.
/finalizethen writes the feature's release note as a changeset, archives the contracts and reasoning, and resets the state directory — all committed to the same branch./releaselater assembles those notes into the changelog; that is release administration, not feature work. - Installer. An
npxCLI that installs into Claude Code (.claude/) and/or Codex (.agents/skills/) from one source of truth, with the pipeline section written once intoAGENTS.md. It updates safely: files your team edited are never overwritten.
Pipeline
/scaffold ─► 00a one real vertical slice, canon layout, test-first (empty repos)
│ Gate A
/bootstrap ─► 00 architecture-context.md: topology, canon mapping, rules
│ Gate B
/intake ─► 01 requirements, business language only
│ Gate 0
/impact ─► 01b predicted blast radius → risk class L0–L3
│
/contract
┌─────────────┼─────────────┐
02a deps 02b backend 02c ui ◄── parallel
└─────────────┼─────────────┘
02d interface seam [FROZEN] page-contract | http-api | graphql | export | import
│ Gate 1
/sequence
┌──────┴──────┐
03a backend 03b ui ◄── parallel; tests named first
└──────┬──────┘
/implement (00b for new modules, 04a, 04b — one task, one commit, TDD loop)
│ Gate 2 every commit
05 CI (typed boundaries, test-with-change, design detector, contract integrity)
│
/impact --verify ─► actual vs predicted, escapes
│
┌──────┴──────┐
06 qa 07 security ◄── parallel, on the branch;
└──────┬──────┘ 08 perf mandatory at L2+
│ Gate 3 approve-merge
/finalize ─► 09a release note + archive + state reset, committed to the branch
│
merge ◄── nothing runs after this
·
/release ─► 09b at release time: assemble notes, retrospective| Gate | After | Decision | |------|-------|----------| | A | scaffold | Is this the shape all future code should copy? | | B | bootstrap | Is this actually how the repo works, and is the canon mapping right? | | 0 | intake | Are these the right requirements? | | 1 | contracts + seam | Right design? Do the halves agree? Are the named types and tests right? | | 2 | every commit | Is this code right, and did the tests come first? | | 3 | qa + security (+ perf), still on the branch | Is this safe to merge? |
Install
No setup on the install side: the package is public on npm, so npx fetches it
directly. PUBLISHING.md covers releasing it, the versioning
policy, and rolling it out across repositories.
npx fullstack-agentic-flow init # interactive
npx fullstack-agentic-flow init --topology mvc --runtime claude-code,codex --tools all --yes
npx fullstack-agentic-flow init … --run-tools # also run the tools' shell install steps
npx fullstack-agentic-flow doctor # check files, config, tools
npx fullstack-agentic-flow update # refresh pipeline files, keep your choices
npx fullstack-agentic-flow tools # list tools and their install commandsWhat gets written:
| Path | Owner after install |
|------|---------------------|
| .ai-agents/ agents, canon, scripts, toolchain | pipeline — refreshed by update unless your team edited a file (then a .incoming copy is written beside it) |
| .ai-agents/state/current-stage.md, rules.config.json | your team — written once |
| .ai-agents/pipeline.config.json, SETUP.md | regenerated from your choices |
| .claude/commands/, .claude/agents/ | pipeline (Claude Code runtime) |
| .agents/skills/flow-*/SKILL.md | pipeline (Codex runtime; invoke as $flow-intake etc.) |
| AGENTS.md | yours — one delimited section holds the pipeline instructions, read by Codex and by Claude Code |
| CLAUDE.md | yours — a managed section with a single @AGENTS.md import, because Claude Code reads AGENTS.md only when no CLAUDE.md exists. --claude-md none skips it |
| .gitignore | yours — one delimited section is managed |
| .claude/settings.json | yours — enabledPlugins / extraKnownMarketplaces keys merged in |
| contracts/ | split topologies only |
Then: an empty repo gets /scaffold and then /bootstrap. A repo with code
gets /bootstrap. After that, /intake.
Upgrading from v1: run init over the existing install. Files you
customised get a .incoming beside them for you to merge. The installer lists
the v1 files that v2 renamed (02d-api-seam.md → 02d-interface-seam.md,
contract-api.md → contract-seam.md) so you can delete them. Re-run
/bootstrap afterwards: the context needs §1.1 topology and §5.0 canon mapping,
and the pipeline-wide rules in §10.0.
Commands
/scaffold first vertical slice, canon layout, test-first (empty repos)
/bootstrap architecture-context.md (once per repo, again after drift)
/scaffold-module canon skeleton for a new backend module
/intake start a feature
/impact risk class before design; --verify before merge
/contract 02a + 02b + 02c in parallel, then 02d freezes the seam
/seam export / import / diff cross-repo contracts (split topologies)
/sequence 03a + 03b in parallel
/implement next task, TDD loop, one commit (--track, --retry-task)
/qa /security /perf review the branch before merging
/finalize release note, archive, state reset — the last commit on the branch
/release assemble the changelog at release time
/status /resumeIn Codex the same commands are skills: $flow-intake, $flow-impact, and so on.
Development of this package
TypeScript, strict (noUncheckedIndexedAccess, exactOptionalPropertyTypes),
zero runtime dependencies, written test-first with node:test.
npm install
npm test # build + 39 tests (installer, adapters, rule scripts, payload vocabulary).gitlab-ci.yml runs the suite on every merge request and publishes on a v*
tag through npm trusted publishing (OIDC) — no NPM_TOKEN, with a provenance
attestation attached automatically. npm version minor && git push
--follow-tags is the whole release.
The rule scripts in payload/ai-agents/scripts/ are plain .mjs with
// @ts-check, so target repos run them without a build step, and this repo
type-checks them with the rest.
Why the two tracks are parallel
Sequential full-stack work has a familiar shape: backend gets built, frontend starts, frontend discovers the API is missing a field, backend changes, frontend adapts. The rework is not a discipline problem. It is structural — the frontend's requirements are discovered while building it, which is after the backend was specified.
So both contracts are written simultaneously, each declaring its side of the API, and a reconciliation agent diffs them before either implementer starts.
That agent — 02d, the interface seam — is where this design earns its keep. It finds the type mismatches, the missing display fields, the disagreement about what an empty result looks like, the error case one side renders and the other never returns. Those are the integration bugs, and it finds them during design rather than during integration.
Once frozen, contract-seam.md outranks both contracts. The UI builds against
mocks that match it exactly. Neither implementer may quietly adapt — a change
comes back to the seam and both tracks are notified.
What the UI track actually specifies
Interface work is where most agentic pipelines go thin — "build the page" and hope. This one specifies:
- Screen structure as a region hierarchy, and visual hierarchy tied to the primary task from intake
- Composition — every element mapped to a component from the approved inventory. Inventing one is a Gate 1 decision, never an implementation detail.
- Every required state — initial load, background refetch, empty-no-data, empty-no-results (a different state, with different copy), partial failure, recoverable error, fatal error, success feedback, in-flight submission. With the actual copy, not a note that copy is needed.
- Forms — field order, where server errors land, double-submit prevention, unsaved-changes behaviour
- Interaction — keyboard paths, focus management on every transition, debounce intervals, motion and its reduced-motion fallback
- Responsive — per breakpoint, in the direction the usage context demands
- Accessibility — against a stated target, tested individually by the QA agent
- Tokens, never literal values — a hex code in a diff is a review failure
The UI implementer's self-check enforces all of it before committing.
Changing the design system
Adopting a UI kit, replacing a component library, or migrating a styling approach
is not a feature. Running it through /intake is a category error: the
feature pipeline assumes the conventions are stable and the feature conforms to
them, and here the conventions themselves are the deliverable.
First, decide which of two things you actually mean.
Selective borrowing — you want a datatable, a chart wrapper, two or three components from a kit. This needs no special handling at all. It is a Gate 1 component approval, which is the mechanism already built in: the UI contract proposes it under §9, a human approves, it enters §8.3, features compose from it.
Wholesale adoption — the kit becomes the design system. Its tokens replace yours, its components become the inventory, existing screens migrate. This is a migration, and it wants its own sequence:
Do it on its own branch, outside the feature pipeline. Integrate the kit, wire its tokens into wherever §8.4 points, and rebuild one real screen with it end to end — including every required state. One screen, not a demo page, because that screen becomes the new exemplar.
Re-run
/bootstrap. This is the step people skip, and skipping it is what produces permanently hybrid codebases. The context does not update itself: until §8.3, §8.4 and §8.10 are re-derived, the UI implementer keeps composing from the old inventory and every new screen is built in the outgoing style.Bootstrap diffs its findings against the existing context and asks before overwriting anything you hand-corrected, so this is safe to run mid-migration.
Check that §8.10 is now a kit-built screen. If the exemplar is still a pre-migration file, every future UI task imitates the thing you are trying to leave.
Expect §11 to grow. Bootstrap classifies old-pattern-plus-new-pattern as a migration in progress: it documents the new one and records the old one as known debt. Give it an explicit rule — "screens under {path} are pre-migration; migrate on touch, do not extend" — so agents neither imitate the old style nor opportunistically rewrite screens the feature did not ask them to.
Two things worth knowing before adopting any large kit:
Put what you use in the inventory, not what the kit ships. A §8.3 listing two hundred components is equivalent to no inventory — the agent cannot choose, so it picks plausibly, and "plausibly" is how you end up with three different modals. Add entries as features genuinely need them.
Record where the kit fights the framework. Kits that ship their own DOM behaviour, plugin initialisation, or jQuery-era widgets frequently conflict with reactive rendering. Wherever you find a workaround, it belongs in §8 as a convention or in §11 as debt. Otherwise every agent rediscovers the same conflict and invents a different workaround for it.
Design decisions worth knowing
One task, one commit, then stop. Not a throughput limitation. It is what keeps review tractable and what makes a wrong turn cost one commit.
Agents stop rather than improvise. When a contract can't be implemented as written, the implementer reports options and waits. An agent that works around a bad contract produces code nobody specified and nobody reviews against anything.
The exemplars do the heavy lifting. architecture-context.md §5.8 and §8.10
contain real, currently-committed files pasted verbatim. Structural imitation
beats description, and it is why generated code matches your house style rather
than the model's defaults.
CI is the backbone. Agents stay light because mechanical checking happens every push without anyone remembering to ask. Only rules that won't false-positive belong there — three reliable checks beat fifteen approximate ones, because a noisy job gets disabled within a week and never comes back.
Judge against stated scale. The architecture context records real volumes. Contracts design to them, the performance agent judges against them. Without that anchor, every review becomes a list of improvements nobody acts on.
Archive the reasoning. Before the merge, contracts and the decisions log are archived into the branch. In a year, when someone asks why a field works the way it does, that is the only record — the diff shows what, never why.
Name the shape, even when it costs a class. Arrays are the fastest way to write the first version and the slowest way to change the tenth. The canon trades a little typing now for a codebase an agent, or a colleague, can read from signatures alone.
Measure before deciding how careful to be. A copy change and a change to how invoices are totalled should not get the same process. The impact agent makes that difference explicit, and the risk class makes it binding.
