@agents-forge/aiqa
v2.0.1
Published
AI Assisted Quality Engineering — single unified agent with playwright-cli snapshot-aware testing. analyst → qa-planner → qa-engineer → qa-reviewer → qa-reporter
Maintainers
Readme
🤖 AIQA — AI Assisted Quality Engineering
One command. Full AI-assisted QA suite — requirements, test plan, test cases, Playwright scripts, a selector-validated review, and an execution report. Run it autonomously in one shot, or step-by-step in your terminal, approving each step as you go.
npx @agents-forge/aiqa https://www.yourwebsite.comWhat makes this different
AIQA is a Super Agent — a single Claude session that orchestrates five specialist subagents, not a chain of disconnected tools. The Super Agent holds the whole pipeline in one context, so insights and snapshots flow in memory from one subagent to the next:
┌──────────────────────────────────────────────┐
│ AIQA — Super Agent │
│ (one Claude session, shared memory) │
└──────────────────────────────────────────────┘
│
analyst → qa-planner → qa-engineer → qa-reviewer → qa-reporter
└──────────────────── subagents ───────────────────────────┘The subagents
| Subagent | Role | Output |
|---|---|---|
| analyst | Explores the site with playwright-cli, captures accessibility snapshots of every page | requirements.md |
| qa-planner | Uses snapshots to assess real UI complexity and assign risk-based priorities | test_plan.md |
| qa-engineer | Generates test cases + Page Objects + Playwright scripts using actual element refs — no guessing | test-cases/, test-scripts/ |
| qa-reviewer | Validates every selector in the specs against the captured snapshots | review_report.md |
| qa-reporter | Runs the suite, captures failure snapshots, writes the stakeholder report | summary_report.md |
Every subagent works against the same playwright-cli accessibility snapshots the analyst captured — so the test scripts reference real DOM elements, and the reviewer can flag any selector that doesn't exist in the ground-truth snapshots.
📐 For a deep dive into how the Super Agent orchestrates the subagents, the runtime flow, and the key design decisions, see docs/architecture.md.
Install
npm install @agents-forge/aiqa
npm install --save-dev @playwright/test
npm install -g @playwright/cli@latest
playwright-cli install --skills agents
npx playwright install chromium
# Optional — enables real accessibility (WCAG) test generation, see below
npm install --save-dev @axe-core/playwrightAuthentication
Auto-detected, and checked before anything else runs — a bad or missing credential fails in under a second instead of after a wasted pipeline run.
| Provider | Setup |
|---|---|
| Anthropic (default) | ANTHROPIC_API_KEY=sk-ant-... in .env, or claude login |
| GitHub Copilot | gh auth login (or GITHUB_TOKEN env, or copilot.token in aiqa.config.ts) — requires a GitHub Copilot license |
Anthropic is preferred when both are available. Force one explicitly with
--provider anthropic / --provider copilot-sdk. See Providers below for details.
Usage
# Full pipeline. In an interactive terminal this runs STEP-BY-STEP by default:
# each step pauses (showing how long it took) and waits for you to press Enter
# before the next one starts. The slow qa-reviewer step is deferred (see below).
npx @agents-forge/aiqa https://my-app.com
# Autonomous — the classic one-shot run, no pauses, all five subagents incl. reviewer.
# (Also the automatic behavior for non-TTY / CI / piped / programmatic / --ui runs.)
npx @agents-forge/aiqa https://my-app.com --auto
# Run the deferred (slow) reviewer later, against the existing run's session (no URL).
# Any step name works as the first arg: analyst | qa-planner | qa-engineer |
# qa-reviewer | qa-reporter.
npx @agents-forge/aiqa qa-reviewer
# Smoke tests only
npx @agents-forge/aiqa https://my-app.com --grep @smoke
# Regression suite to custom dir
npx @agents-forge/aiqa https://my-app.com --grep @regression --dir ./aiqa-output
# Skip the review subagent
npx @agents-forge/aiqa https://my-app.com --skip qa-reviewer
# Resume after a crash
npx @agents-forge/aiqa https://my-app.com --resume
# Interactive — no URL given, so this opens a local browser page to configure the run
npx @agents-forge/aiqa
# Interactive, terminal-only — for headless/remote/no-browser machines
npx @agents-forge/aiqa --no-ui
# Force the browser UI even when a URL was already passed (pre-fills it into the form)
npx @agents-forge/aiqa https://my-app.com --ui
# Unattended / CI — don't prompt before overwriting playwright.config.ts
npx @agents-forge/aiqa https://my-app.com --force
# Inside VS Code — open the generated docs in Markdown preview when done
npx @agents-forge/aiqa https://my-app.com --open
# Run on a GitHub Copilot license instead of Anthropic
npx @agents-forge/aiqa https://my-app.com --provider copilot-sdk
# Verbose mode — full raw tool-call detail instead of the clean step view
npx @agents-forge/aiqa https://my-app.com --verbose
# Watch the Analyst explore the site in a real, visible browser window
npx @agents-forge/aiqa https://my-app.com --headed
# Terminal only — don't open the live browser dashboard
npx @agents-forge/aiqa https://my-app.com --no-dashboard💡 Using the Claude Code VS Code extension? Run
npx @agents-forge/aiqa --init-skillonce per project to write.claude/skills/aiqa/SKILL.md— after that,/aiqa <url> [flags]runs the pipeline directly from the chat panel. This one-time step is needed because that file isn't part of the npm package (npm installalone won't create it).
Step-by-step mode (default in a terminal)
When you run AIQA in a real interactive terminal, it now walks the pipeline one step at a time:
- It runs a step (e.g. the analyst), then pauses, printing how long that step took.
- You review the freshly-generated artifact (
requirements.md, thentest_plan.md, …). - Press Enter to continue to the next step (or Ctrl+C to stop and pick up later). The live dashboard shows the same pause with a Continue button — either one works.
At the end it reports the total active time in minutes — the time you spend reading between steps is deliberately not counted (only the model's actual working time is).
The slow qa-reviewer step is deferred. It's skipped from the interactive chain so you're not
blocked on it. Run it whenever you're ready, against the same session:
npx @agents-forge/aiqa qa-revieweraiqa <step> runs any single step (analyst, qa-planner, qa-engineer, qa-reviewer,
qa-reporter) against the existing on-disk session — no URL needed (it's read from session.json).
Prefer the classic one-shot run? Pass --auto to run everything autonomously with no pauses.
Step-by-step also turns itself off automatically when there's no interactive terminal — CI, piped
stdin, programmatic runAIQA() calls, and the --ui browser flow all stay autonomous.
Live dashboard
Every interactive run now opens a live dashboard in your browser. The terminal output is exactly the same as before; the dashboard is a second view of the same run.
- From the
--uiform: after you click Run AIQA, the same page turns into the dashboard. - From a URL run (
npx @agents-forge/aiqa https://my-app.com): the dashboard opens on its own. This coversaiqa <step>runs too.
What it shows:
| While running | When it pauses (step-by-step) | When it's done |
|---|---|---|
| The five steps (done, running, waiting) with a live timer and tokens used so far, plus counts of files written, pages captured and test cases scripted. A live activity feed shows files, agent notes, snapshot reviews and tasks, and you can filter it to Files or Agent notes. | A Continue to <next step> button that works like pressing Enter, what the step produced (with a Preview of requirements.md / test_plan.md), and the aiqa qa-reviewer command for the deferred reviewer. | The result, total time, passed/failed tests (from Playwright's results.json), files created and tokens used (with how many came from the prompt cache). Also time and tokens per step, the output files, any failing tests, and next-step commands with Copy buttons. |
If the run finds an existing playwright.config.ts, the dashboard asks Keep my config / Allow
overwrite too. Whichever answers first wins, the dashboard or the terminal's y/N.
When the run ends, a self-contained copy is saved to .aiqa/reports/dashboard.html. You can reopen
it any time or attach it to a PR, even after the CLI has exited. It has a light/dark toggle, which
starts from your OS setting, and it works at phone width.
Turning it off: pass --no-dashboard, or set dashboard: false in aiqa.config.ts. --no-ui
turns it off too. It never opens for non-TTY runs (CI or piped). Over SSH, or on Linux with no
display, AIQA only prints the URL instead of trying to open a browser.
Security: the dashboard server listens on 127.0.0.1 only, behind a random per-run URL token.
It rejects requests with an unexpected Host header, which blocks DNS-rebinding, and it stops when
the CLI exits.
Progress output extras
- Per-step timing: each step prints its own elapsed time when it
finishes (
✅ analyst completed in 42s), not just the total run time at the end. - Test case count: right after
qa-engineerfinishes, a code-computed line shows how many test cases were drafted (test-cases/*.md) vs actually scripted (test-scripts/*.spec.ts) — a quick pulse check, with a⚠️line if any drafted case didn't get scripted. (qa-reviewer's own, more thorough traceability check still runs separately and shows up inreview_report.md.) - Clickable preview links:
requirements.mdandtest_plan.mdget a clickable link right after they're written, opening in VS Code's rendered Markdown preview — only shown when running in a terminal that supports clickable links (VS Code's integrated terminal, Windows Terminal) and thecodeCLI is on PATH; silently omitted elsewhere (e.g. classiccmd.exe) rather than showing garbled escape characters. This sets up the same*.md → Markdown previewassociation in.vscode/settings.jsonthat--openalready uses — automatically, the first time a link is shown, even if you never pass--open.
Bring your own requirements (optional)
You can seed the pipeline with an existing requirements document — the analyst still explores the site and then merges the two, rather than skipping exploration:
- Drop your requirements
.mdor.pdffile into.aiqa/requirements/(the folder is created for you on first run) — or upload it through the--uiform instead. - Run AIQA normally (or just
npx @agents-forge/aiqafor the browser UI / prompts). - The analyst explores the URL, then writes
.aiqa/requirements/requirements.mdwith two clearly-labelled parts:- Part A — Existing Requirements (carried over from your doc, tagged
[existing]) - Part B — Newly Discovered Requirements (the addon found by exploring, tagged
[discovered])
- Part A — Existing Requirements (carried over from your doc, tagged
Nothing from your document is dropped, and the merge happens automatically — no mid-run prompt.
Naming: your doc can have any name. If you happen to name it
requirements.md(the same as the generated output), AIQA preserves your copy asexisting-requirements.mdfirst, then writes the merged result torequirements.md— so your original is never lost.
Test scope & speed
You choose which tests run. The QA Engineer always writes the full suite. The test scope only decides which tagged tests the QA Reporter runs:
- In a terminal, if you didn't pass
--grep, AIQA asks before the run: 1) Smoke only (fastest; this is what Enter picks), 2) Full suite, or 3) a custom tag. - In the
--uiform, picking a test scope is required. Run AIQA stays disabled until you tick at least one test type or Full suite. - On the command line, use
--grep <tag>for one tag (e.g.--grep @regression), or--grep ""for the full suite with no filter. You can also setgrepinaiqa.config.ts. - Programmatically, pass
grep: ""torunAIQA()for the full suite, orgrep: "@tag"for one tag. - When nobody can be asked (CI, piped stdin, programmatic
grep: undefined, the/aiqaslash command), AIQA falls back to smoke only and prints🏷️ Tests: @smoke (default …), so the choice is always visible.
Two fast-feedback defaults keep an automatic run quick, and both are easy to override:
- Smoke only when nobody chose. This is the fallback described above.
- Chromium only, by default (via your machine's already-installed Chrome
—
channel: "chrome"in the generated config, so there's no extra browser download). The generatedplaywright.config.tshasfirefox/webkit/mobileprojects commented out, ready to enable:
Uncomment what you want, then// { name: "firefox", use: { ...devices["Desktop Firefox"] } }, // { name: "webkit", use: { ...devices["Desktop Safari"] } }, // { name: "mobile", use: { ...devices["Pixel 5"] } },npx playwright testpicks them up.
Together these mean a typical automatic run is testing a fraction of the full matrix (one browser × smoke tests) — intentionally, so you get fast feedback first and opt into the full regression/cross-browser pass when you're ready for it.
Accessibility testing
Real automated WCAG auditing via @axe-core/playwright
(Deque Systems' engine — the standard tool for this in the Playwright ecosystem),
not a weaker proxy based on ARIA structure alone:
npm install --save-dev @axe-core/playwright- Two conditions gate generation — both required. AIQA checks whether
@axe-core/playwrightis installed (same probe-first pattern already used forplaywright-cli), and whether you actually selected Accessibility for this run (--ui's checkbox,--grep @accessibility, or a custom tag — see below). Installing the package alone does not trigger generation; it's opt-in every run, not automatic just because it's present. If either condition isn't met, that step is skipped with a one-line explanation (which one wasn't met, and how to fix it) — nothing else in the run is affected either way. Note: "Full suite" alone does not select Accessibility — it's treated as a distinct, heavier category you opt into explicitly, not folded into "run everything." - If
@axe-core/playwrightisn't installed and Accessibility was selected, AIQA attempts an automatic install —npm install --save-dev @axe-core/playwrightin your project — before falling back to a manual install hint if that fails. Lower-risk than k6's system-level install below (no sudo, no system package manager — just a devDependency in your own project). The attempt only ever fires when Accessibility was actually selected and the package is missing; it never runs unconditionally, and it never blocks the run on failure. - Scope: WCAG 2.1 AA (via axe's
wcag2a/wcag2aa/wcag21a/wcag21aarule tags — the same four tags Playwright's own accessibility-testing docs use to match "WCAG A and AA" coverage), matching the targetqa-planneralready documents in the test plan. Not configurable yet — see Configuration for what is adjustable today. - Only critical/serious violations fail the test. Moderate/minor
violations are still captured (written to
.aiqa/reports/accessibility/accessibility-<page>.json, kept separate from Playwright's owntest-results/) and surfaced insummary_report.md's "Accessibility Findings" section as warnings — they just don't fail the run. Most real sites carry some pre-existing minor issues; failing on every single one would make nearly every run "fail" and just train people to ignore the results. - The generated test file lives in its own subfolder —
.aiqa/test-scripts/accessibility/accessibility.spec.ts— separate from per-module specs, but still inside the same PlaywrighttestDir(its default recursive discovery already covers subfolders, so no config changes are needed for it to actually run). - Runs like any other tagged category — no special flag needed. Select
it via the interactive prompt's custom-tag option,
--grep @accessibility, or the--uiform's Accessibility checkbox (Non-Functional section). - Full scan results are attached to the Playwright HTML report, not just
violations — view them via
npx playwright show-report(each accessibility test has an "accessibility-scan-results" attachment). Useful for debugging why a rule didn't fire the way you expected, since axe's passes/inconclusive data is included alongside violations.
Performance testing
Real load testing via k6 (Grafana's open-source, AGPL-3
licensed load-testing tool) — a separate CLI binary, not an npm package,
and not a Playwright test. It runs via its own k6 run command, entirely
outside npx playwright test.
# Manual install (AIQA also attempts this automatically — see below):
winget install k6 --source winget # Windows
brew install k6 # macOS
# Linux: https://k6.io/docs/get-started/installation/#linux- Two conditions gate generation — both required, same rule as
Accessibility: k6 must be available, and you must have actually
selected Performance for this run (
--ui's checkbox,--grep @performance, or a custom tag). Installing k6 alone does not trigger generation. "Full suite" alone does not select Performance either — same distinct, explicit opt-in category as Accessibility. - If k6 isn't found and Performance was selected, AIQA attempts an
automatic install —
wingeton Windows,brewon macOS, the official apt/dnf repo on Linux — before falling back to a manual install link if that fails or isn't supported on your platform. This is a deliberate exception to how every other external tool is handled here (playwright-cliis never auto-installed, only probed with a manual hint) — made because k6's per-OS install commands are simple and well-documented. The attempt only ever fires when Performance was actually selected and k6 is missing; it never runs unconditionally, and it never blocks the run on failure. - Scaled from two inputs — concurrent users and sustained duration —
not a full k6 configuration wizard. In
--ui, these two number fields (default 10 users, 2 minutes) only appear once the Performance checkbox is ticked; via CLI use--performance-users <n>/--performance-duration <n>(minutes); viaaiqa.config.ts, setperformance: { users, durationMinutes }. Everything else is derived: stress uses 2× the users (a conservative "probe beyond capacity" multiplier) over half the sustained duration (rounded up) so total run time stays bounded; spike uses a baseline rate of users÷10/s bursting to users/s, in fixed short 10s stages regardless of scale (spikes don't scale with sustained duration). Ramp shape and thresholds stay fixed — hand-edit the generated script for anything more advanced. - Three sequenced scenarios in the generated
.aiqa/performance/load-test.js(never run simultaneously — each starts only after the previous one's total duration elapses, via k6'sstartTime, so they don't skew each other's metrics):- load —
ramping-vus: ramp up → steady → ramp down. Standard baseline. - stress —
ramping-vusat a higher target than load, to probe beyond normal capacity. - spike —
ramping-arrival-rate(deliberately notramping-vus) — a spike test is fundamentally about a sudden burst of request rate arriving, independent of response time, which is what arrival-rate executors model.
- load —
- k6's own threshold mechanism is the pass/fail signal — default
thresholds (
http_req_duration: p(95)<500ms,http_req_failed: <1%) are baked into the generated script; k6 itself exits non-zero if any is breached. No separate severity system is invented on top, unlike Accessibility's critical/serious split — k6 already has its own answer. - Reporting stays in k6's own native format, not reproduced inline.
performance-report.html(a standalone, browser-viewable report generated by the k6-reporter library, pinned to a specific released version — notmain— to avoid unpinned remote code changing under you) plusperformance-summary.json(machine- readable) land in their own.aiqa/reports/performance/folder — kept separate from Playwright's owntest-results/, since these aren't Playwright-native artifacts.summary_report.mdonly gets a short executive rollup (pass/fail + key metrics) that links to the HTML report — exactly how it already treats Playwright's own report today (referenced, never inlined). - Beyond users/duration, script scope is still fixed — ramp shape, the stress/spike multipliers, and thresholds aren't independently configurable; edit the generated file directly if you need different numbers.
Visual setup — --ui
Running with no URL now opens a local page in your default browser to configure the run — no more terminal Q&A by default:
npx @agents-forge/aiqaPass --no-ui for the old terminal-prompt flow instead (recommended for
headless/remote/no-browser machines — see Troubleshooting).
--ui can also be passed explicitly to force the browser page even when a
URL was already given on the command line (it pre-fills the form).
The form covers everything the terminal prompts ask, plus a file upload:
- Target URL (pre-filled if you also passed one on the command line), with a "Show the browser during site exploration" checkbox — see Watching the Analyst explore below
- Provider — Anthropic (Claude) or GitHub Copilot. Required — you must
pick one (there's no auto-detect in the form); "Run AIQA" stays disabled
until you do. GitHub Copilot shows an inline note that it needs the Copilot
CLI installed and authenticated (
gh auth login) plus a Copilot license, and reveals an optional model box — leave it blank to use your Copilot account's default model, or type a model id to pin one. - Existing requirements (optional) — upload a
.mdor.pdffile directly in the form (.docxisn't supported yet — see below), instead of manually dropping it into.aiqa/requirements/first. Uploads are capped at 20MB and the filename is sanitized before being saved. - Test Types (data-driven, so this list grows over time). Required: Run AIQA stays
disabled until you tick at least one, or Full suite.
- Functional — Smoke, Regression, Sanity
- Non-Functional — Accessibility (real WCAG 2.1 AA auditing via
@axe-core/playwrightwhen installed) and Performance (real k6 load testing via k6 when available — AIQA will try to auto-install it if it's missing and this is selected). Ticking Performance reveals two more fields — concurrent users and sustained duration — that scale the generated scenarios. - Full suite — overrides the Functional checkboxes above, runs smoke + regression + negative + edge. Does not include Accessibility or Performance — both are distinct, explicit opt-ins (see their sections above), not folded into "everything."
- ⚠️ Sanity is still a filter placeholder — qa-engineer doesn't
generate
@sanity-tagged tests yet, so selecting it currently matches no tests. It's there so the option is visible and reachable once real generation for that category ships.
The server is loopback-only (127.0.0.1, an OS-assigned free port, behind
a random per-run URL token), so it's never exposed to your network. After you
submit, the same page becomes the live dashboard for the
run, and the server stops when the CLI exits. With --no-dashboard it shuts down
as soon as you submit, as before. The form's answers are authoritative — they override any --provider
or --grep also passed on the command line. There's no per-agent opt-out in
the form — all five subagents always run via --ui, so --skip is
ignored too when combined with it. If you need to skip a specific agent, run
without --ui (or with --no-ui) and pass --skip on the command line.
Requirements doc formats today:
| Format | Status |
|---|---|
| .md | Supported — read as-is |
| .pdf | Supported — Claude's Read tool reads PDFs natively during the run, no extra parsing needed |
| .docx | Not yet — would need a text-extraction dependency (e.g. mammoth) not currently included |
Watching the Analyst explore
playwright-cli (which the Analyst uses to explore your site) is headless
by default — no visible browser, nothing to watch. Pass --headed on the
CLI, or tick "Show the browser during site exploration" in --ui, to
have it open a real, visible browser window instead.
- Scoped to the Analyst's exploration step only (Step 1).
qa-engineer's per-module test generation andqa-reporter's failure-snapshot capture stay headless regardless — and the actual automated test run (npx playwright test) always stays headless too, for CI-safety and so parallel workers don't fight over screen focus. - Needs a real display. Fine on your own machine; leave it off (the
default) on headless CI runners or remote servers with no display attached
—
playwright-cli open --headedwill simply fail there.
⚠️ Installed into an existing project?
AIQA writes everything it generates under .aiqa/. The only file placed at your repo
root is playwright.config.ts (so npx playwright test can auto-discover it). Since
it's typically run inside an existing project, the Super Agent guards that one file:
- If a
playwright.config.tsalready exists, AIQA asks once, before the run, whether it may overwrite it. - Decline, and your existing config is left untouched — the pipeline runs against it as-is.
- In a non-interactive shell (CI, piped) it never silently overwrites — pass
--force(orforce: true) to opt in. - Use
--dir <path>to run the whole pipeline (and its.aiqa/folder) in a different directory.
Providers
AIQA runs on a pluggable engine — the pipeline itself (prompts, session, output layout) is identical regardless of which one executes it:
| Provider | Engine | Requires |
|---|---|---|
| anthropic (default) | Claude Agent SDK | ANTHROPIC_API_KEY or claude login |
| copilot-sdk | @github/copilot-sdk (optional dependency) | The @github/copilot-sdk package, plus a GitHub Copilot license authenticated via gh auth login |
@github/copilot-sdk is an optional dependency — a normal
npm install @agents-forge/aiqa will pull it if your environment allows,
but a Claude-only install that skips it works fine: the Copilot engine is
loaded lazily, only when that provider is actually selected. If it's missing
and you select copilot-sdk, AIQA fails fast with an install hint
(npm install @github/copilot-sdk) rather than a crash — the Claude path is
never affected.
Select one explicitly with --provider <name> or provider: "..." in
aiqa.config.ts. On the CLI/programmatic path, omitting it auto-detects
(Anthropic wins if both are available); the --ui form has no
auto-detect — it requires you to pick a provider before the run can start.
Choosing the Copilot model. It's optional — omit it to use your Copilot
account's own default model. To pin a specific one, set it any of three
ways (CLI flag > aiqa.config.ts > the --ui field): --copilot-model <id>,
copilot.model below, or the model box that appears when you pick GitHub
Copilot in the --ui form (leave that box blank for the account default).
Run node scripts/spike-copilot-sdk.mjs to list what your plan entitles.
// aiqa.config.ts
export default defineConfig({
provider: "copilot-sdk",
copilot: {
model: "claude-sonnet-4.5", // OPTIONAL — omit to use your Copilot account's
// default. Or e.g. "gpt-5"; run the spike script
// (below) to list what your plan has access to
token: undefined, // optional — omit to use `gh auth login` / GITHUB_TOKEN
timeoutMs: 60 * 60 * 1000, // how long to wait for a response (not a cost/turn cap)
},
});Before relying on copilot-sdk for real work, run the included spike
script once to confirm your setup end-to-end (auth, model access, and the
real tool names/argument shapes the Copilot CLI runtime uses on your
machine/version):
node scripts/spike-copilot-sdk.mjs
node scripts/spike-copilot-sdk.mjs --model claude-sonnet-4.5Known v1 trade-offs of the copilot-sdk engine:
- AIQA imposes no model of its own — omit
copilot.modeland the run uses your Copilot account's default; pin one explicitly (via--copilot-model, config, or the--uifield) if you want a specific model. Catalog and plan entitlements vary, so the spike script lists what's available to you. - Both engines share identical subagent prompts today; a model reached
through Copilot other than Claude may be less reliable at the multi-step
tool-chaining the prompts assume. A Claude model selected via Copilot is
expected to behave closest to the default
anthropicpath. - Cost tracking (
totalCostUsd) isn't available on this engine yet.
Note for library consumers:
detectProvider()is nowasyncand returns{ provider, reason }instead of a bare string — a breaking change from earlier versions that only ever guessed at auth (it never actually checked login state).
What it produces
Everything lands under .aiqa/ so your repo root stays clean — only
playwright.config.ts is written to the root (so npx playwright test finds it):
.aiqa/
├── session.json ← pipeline state (resumable)
├── auth-state.json ← saved login session (if any)
├── snapshots/
│ ├── home.yml ← accessibility trees (ground truth)
│ ├── login.yml
│ └── failure-<test>.yml ← captured on test failure
├── requirements/
│ └── requirements.md ← analyst subagent: business analysis
├── test-plan/
│ └── test_plan.md ← qa-planner subagent: strategy + risk analysis
├── test-cases/
│ ├── login.md ← qa-engineer subagent: plain-English test cases
│ └── checkout.md
├── test-scripts/
│ ├── pages/
│ │ ├── login.page.ts ← qa-engineer subagent: Page Object (locators + actions per module)
│ │ └── checkout.page.ts
│ ├── login.spec.ts ← qa-engineer subagent: Playwright scripts, using the page objects above
│ ├── checkout.spec.ts
│ └── accessibility/
│ └── accessibility.spec.ts ← qa-engineer subagent: axe-core WCAG scan (if selected + available)
│ still a real Playwright test — nested here so testDir's
│ default recursive discovery picks it up automatically
├── performance/
│ └── load-test.js ← qa-engineer subagent: k6 script (if selected + available) — NOT a
│ Playwright spec, lives at the top level, its own separate `k6 run`
└── reports/
├── review_report.md ← qa-reviewer subagent: selector validation + quality review
├── summary_report.md ← qa-reporter subagent: stakeholder report
├── dashboard.html ← saved copy of the live dashboard (when it was open)
├── test-results/ ← results.json, results.xml, traces, failure screenshots
├── accessibility/
│ └── accessibility-<page>.json ← per-page WCAG violations (if accessibility testing ran)
├── performance/
│ ├── performance-report.html ← k6-reporter HTML report (if performance testing ran)
│ └── performance-summary.json ← k6 machine-readable summary
└── playwright-report/ ← Playwright HTML report (npx playwright show-report .aiqa/reports/playwright-report)
playwright.config.ts ← multi-browser config at repo root (overwrite-guarded)Programmatic Usage
import { runAIQA } from "@agents-forge/aiqa";
const result = await runAIQA({
target: "https://my-app.com",
grep: "@smoke",
skip: ["qa-reviewer"],
cwd: "./aiqa-output",
resume: false,
force: false,
open: false,
existingRequirements: ".aiqa/requirements/my-reqs.md",
model: "claude-sonnet-4-6",
paths: { reports: "reports" },
provider: "anthropic", // or "copilot-sdk" — omit to auto-detect
copilot: { model: "claude-sonnet-4.5" }, // only used when provider is "copilot-sdk"
performance: { users: 25, durationMinutes: 5 }, // only used when Performance testing is selected
verbose: false,
interactive: false, // true → step-by-step gating (needs a TTY; no-ops otherwise)
// onlyStep: "qa-reviewer", // run ONE step against the existing on-disk session, then return
onProgress: (e) => { // structured progress events, alongside the normal console output
if (e.type === "step-end") console.log(`${e.step}: ${e.summary}`);
},
});
console.log(`Passed: ${result.passed}`);
console.log(`Session: ${result.sessionFile}`);
console.log(`Total: ${result.totalDuration}ms`, result.stepDurations); // per-step ms (interactive/onlyStep)| Option | Meaning |
|---|---|
| target | URL to analyse (required) |
| grep | Three states: omit/undefined → smoke-only fallback (programmatic runs can't be asked); "" (empty string) → explicit full suite, no filter; "@tag" → run only tests matching that tag/pattern. See Test scope & speed. |
| skip | Subagents to skip, by name |
| cwd | Output directory (where .aiqa/ is created) |
| resume | Resume a previous run |
| force | Skip the playwright.config.ts overwrite prompt |
| open | Open the generated docs in VS Code Markdown preview |
| existingRequirements | Path to an existing requirements .md to merge in |
| model | Claude model id (used when provider is anthropic) |
| paths | Override the .aiqa base + subfolder names |
| provider | "anthropic" (default) or "copilot-sdk" — see Providers |
| copilot | { model?, token?, timeoutMs? } — only used when provider is "copilot-sdk" |
| performance | { users?, durationMinutes? } — scales the generated k6 scenarios; only used when Performance testing is selected. See Performance testing. |
| verbose | Show tool-call details |
| interactive | Step-by-step gating — run each step as its own session and pause for Enter between them (excludes qa-reviewer). No-ops without a TTY. The CLI enables this by default in a terminal; see Step-by-step mode. |
| onlyStep | Run exactly one step (analyst…qa-reporter) against the existing on-disk session, then return. Powers aiqa <step>. |
| onProgress | (event: ProgressEvent) => void. Receives structured events (run-start, step-start, step-end, file, narration, snapshots, task, tokens, notice, gate, question, run-end, …) next to the console output. This is what powers the live dashboard; use it for your own UI, CI annotations or notifications. A listener that throws never breaks the run. The ProgressEvent type is exported. |
| gateSignal | () => Promise<void>. Step-by-step only: when the returned promise resolves first, the run continues as if Enter had been pressed. |
| answerSignal | (questionId) => Promise<string>. Answers a yes/no prompt (currently the playwright.config.ts overwrite question), racing the terminal. |
result.stepDurations is { [step]: activeMs } for each step that ran (interactive/onlyStep runs), and
result.totalDuration is wall-clock for autonomous runs or the sum of active step time (excluding the
Enter-gate waits) for interactive/onlyStep runs.
Configuration — aiqa.config.ts
Drop an aiqa.config.ts (or .mjs / .js / .json) in your project root to set
defaults. CLI flags override the config file, which overrides the built-in defaults.
Loaded at runtime via jiti — no build step.
import { defineConfig } from "@agents-forge/aiqa";
export default defineConfig({
// Run defaults (any of these is overridden by the matching CLI flag)
url: "https://my-app.com",
grep: "@smoke",
skip: ["qa-reviewer"],
open: true,
force: false,
dashboard: true, // false → never open the live browser dashboard (same as --no-dashboard)
// Model
model: "claude-sonnet-4-6",
// Rename the .aiqa base + subfolders to taste
paths: {
base: ".aiqa",
requirements: "requirements",
plan: "test-plan",
testCases: "test-cases",
testScripts: "test-scripts",
reports: "reports",
performance: "performance",
},
// Provider — see the "Providers" section for details
provider: "anthropic", // or "copilot-sdk"
copilot: { model: "claude-sonnet-4.5" },
// Scales the generated k6 scenarios — only used when Performance testing
// is selected. See "Performance testing" for how stress/spike derive from these.
performance: { users: 25, durationMinutes: 5 },
});defineConfig() is optional but gives you typed autocomplete. Folder renames flow
everywhere automatically — including the generated playwright.config.ts paths.
Skipping subagents
Each stage is a subagent you can skip by name (analyst, qa-planner, qa-engineer, qa-reviewer, qa-reporter):
# Analysis + planning only (no tests)
npx @agents-forge/aiqa https://my-app.com --skip qa-engineer,qa-reviewer,qa-reporter
# Skip the review subagent
npx @agents-forge/aiqa https://my-app.com --skip qa-reviewer
# Write tests but don't run them (skip the reporter subagent)
npx @agents-forge/aiqa https://my-app.com --skip qa-reporterResumable pipeline
If the Super Agent crashes mid-run:
npx @agents-forge/aiqa https://my-app.com --resumeThe session file at .aiqa/session.json tracks which subagents completed. On
--resume, any subagent whose output already exists on disk is marked done and
skipped, so the pipeline picks up where it left off.
Troubleshooting
| Error | Fix |
|---|---|
| playwright-cli not found | npm install -g @playwright/cli@latest |
| Browser not installed | npx playwright install chromium |
| Authentication failed (Anthropic) | Set ANTHROPIC_API_KEY or run claude login |
| Authentication failed (GitHub Copilot) | Run gh auth login, or set GITHUB_TOKEN / copilot.token. Requires a Copilot license |
| Pipeline crashed | Re-run with --resume |
| Won't overwrite my config in CI | Pass --force (non-interactive runs never overwrite playwright.config.ts silently) |
| Want to preview docs in VS Code | Run inside VS Code with --open (opens requirements / test plan / summary in Markdown preview) |
| Want to run tests interactively | npx playwright test --ui |
| Run "looks stuck" with no output for a while | Expected during long steps — a heartbeat line prints periodically showing the current step and idle time, so it's not actually silent. If it's genuinely stuck, Ctrl+C then re-run with --resume. |
| qa-reporter taking a long time | It only runs @smoke-tagged tests on Chromium by default — see Test scope & speed. If it's still slow, check how many tests are tagged @smoke and whether the target app itself is slow to load. |
| Running interactively on a headless/remote/SSH machine | Pass --no-ui — the default browser UI won't have anywhere to open. Or just pass the URL/flags directly so no prompt of any kind triggers. |
| Browser UI didn't open automatically | Copy the http://127.0.0.1:<port>/<token>/ URL printed in the terminal and paste it into a browser yourself — the server is already listening and waiting. |
| Don't want the live dashboard | Pass --no-dashboard (or dashboard: false in aiqa.config.ts). The terminal output is unchanged either way. |
| Dashboard says "Lost connection to the CLI" | The CLI stopped (Ctrl+C or a crash) before the run finished. Re-run with --resume. Finished runs keep their final state, and a copy is saved at .aiqa/reports/dashboard.html. |
| Accessibility tests never get generated | Two possible causes, both required to pass: (1) @axe-core/playwright isn't installed — AIQA tries to auto-install it automatically when Accessibility is selected (npm install --save-dev @axe-core/playwright in your project), but install manually if that fails; (2) Accessibility wasn't selected for this run — pick it via --ui, --grep @accessibility, or a custom tag ("Full suite" alone doesn't count). The run prints a 🔧/✅/ℹ️/⚠️ line stating exactly what happened. See Accessibility testing. |
| Performance test never gets generated | Same two-cause pattern as Accessibility: (1) k6 isn't available — AIQA tries to auto-install it automatically when Performance is selected, but this can fail (unsupported platform, no package manager, needs elevated permissions) — install manually from k6.io if so; (2) Performance wasn't selected for this run — pick it via --ui, --grep @performance, or a custom tag. The run prints a 🔧/✅/ℹ️/⚠️ line stating exactly what happened. See Performance testing. |
| k6 auto-install didn't work | On Windows, AIQA installs the exact GrafanaLabs.k6 winget package ID (--id ... --exact) specifically because the bare name k6 is ambiguous in the winget repository (multiple unrelated packages share the substring) and fails silently in non-interactive mode otherwise. Elsewhere it only tries one package manager per OS (brew/apt-or-dnf) and never asks for a password interactively, so it fails fast rather than hanging if elevated permissions are needed. Install k6 yourself from k6.io and re-run. |
| No clickable link for requirements.md/test_plan.md | Expected if your terminal isn't known to support clickable links (only VS Code's integrated terminal and Windows Terminal are detected), or the code CLI isn't on PATH — the files are still written normally, just without a link. See Progress output extras. |
License
MIT
