@seldonqa/cli
v0.11.0
Published
Seldon QA - generate a QA test plan from any codebase and run AI agents against your live app, from your terminal or CI.
Downloads
371
Maintainers
Readme
Seldon QA
Write your QA test plan as code, then run AI agents that use your product like a real user — from your terminal or CI.
seldon keeps a baseline of user-facing test cases in a committed seldon.yml. Point it at a live URL and autonomous agents exercise each case — click, type, navigate, verify — then hand back a clean, actionable report for you or your coding agent to act on.
npm install -g @seldonqa/cliWhy
AI is shipping more code than any team can test by hand. Seldon closes the gap: describe the journeys that matter once, in plain English, and agents run them against every deploy — telling you exactly what broke, why, and what to do about it.
Get started
npm install -g @seldonqa/cli
seldon login seldon_your_api_key # or: seldon config set-key <key>From inside a project (one baseline per project):
seldon init # scaffold a template seldon.yml
seldon cases add --name "Login works" \
--step "go to /login" --step "sign in" --criteria "lands on the dashboard"
seldon validate # check the file locally (the server re-validates on run)
seldon run --url https://staging.example.com
seldon logs --follow # watch the agents test, live
seldon report --md # clean, actionable resultsHighlights
- Test plan as code. Your cases live in a committed
seldon.yml— reviewed in PRs, versioned with your app.seldon initscaffolds it;seldon cases addappends more. - Agents that behave like users.
seldon rundrives your live app and judges each case PASS / FAIL / BLOCKED / TIMED_OUT. - Reports built for action. Every failure carries an
expected/observedpair and a single concrete recommendation —--mdfor pull requests,--jsonfor coding agents, colorized plain text for the terminal. - Watch it think.
seldon logs --followstreams the agent's reasoning, framed by test case and step. - Interrupt-safe. A submitted run executes server-side. Press Ctrl-C and the CLI detaches without stopping the run — re-attach any time with
seldon report. - CI-native gate.
seldon run --fail-on failexits non-zero so a red QA run can block a merge.
Configuration — seldon.yml
Your baseline is a single committed file at the project root — the plain-English cases themselves, not id overrides:
name: web-app # optional lineage name; defaults to the git remote owner/repo
labels: [nightly, team-web] # optional run tags for profiling; merged with `--label`
run:
url: https://staging.example.com # default target; override per run with --url
username: [email protected] # optional login the agent signs in with
password: hunter2 # optional; prefer a managed secret in prod
testCases:
- name: Login works
id: login-works # optional stable id; defaults to a slug of the name
url: /login # optional per-case path (relative) or absolute URL override
steps:
- go to /login
- enter valid credentials and submit
criteria: the user is signed in and lands on the dashboardReading a report
The report lists every case worst-first, then expands each failure:
Seldon QA — Checkout & account flows · 2/6 passed (33%) · ❌2 · 🚫1 · ⌛1
FAIL BLOCKER TC-201 Complete payment
BLOCKED TC-301 View order history
PASS TC-101 Add item to cart
TC-201 Complete payment
acceptance: A valid card places the order and shows a confirmation.
expected: Order is placed and a confirmation screen appears.
observed: The Pay button spins indefinitely; no order is created.
recommendation: Wire the Pay button's submit handler to /orders and surface a failure toast.
repro: Add an item, open /checkout, enter a valid card, click Pay- Verdicts:
PASS·FAIL(a real defect, with aseverity) ·BLOCKED(couldn't be evaluated — e.g. a missing precondition) ·TIMED_OUT. - recommendation appears on
FAIL/BLOCKEDcases: the single strongest next step to get the case passing. --fullexpands every failing case (the terminal view caps at 5 by default);--jsonemits the same data as a stable object.
In CI
seldon run --url "$PREVIEW_URL" --fail-on fail--fail-on fail returns a non-zero exit code when any case fails; --fail-on any also gates on blocked or timed-out cases. Non-interactive shells block until the run finishes; --no-wait submits and returns immediately (re-read later with seldon report). With --json the report is stable, machine-readable output you can post to a pull request or hand a coding agent for an automated fix loop.
Local targets — automatic tunnels
Point --url at a local address — localhost, 127.0.0.1, 0.0.0.0, ::1, or a bare :PORT — and Seldon exposes it to the QA agent through a temporary, authenticated Cloudflare tunnel for the duration of the run. No deploy, no manual tunnelling:
seldon run --url http://localhost:3000You'll see a one-line warning that the local address will be exposed and a [Y/n] prompt (defaults to yes). The public hostname is protected by Cloudflare Access, so only this run's agent can reach it, and the tunnel is torn down automatically when the run finishes or you press Ctrl-C. In CI or any non-interactive shell the prompt is skipped automatically; pass --auto-approve to skip it explicitly:
seldon run --url localhost:3000 --auto-approve # e.g. in a CI jobRemote --urls are used as-is (no tunnel).
Tunnels are short-lived. Each tunnel has a hard maximum lifetime (20 minutes by default) and is reclaimed by the server once it's exceeded, regardless of the run's state — so the run must finish within that window. When a tunnel comes up the CLI prints the exact cap for your account (… active for up to 20 min; the run must finish within that window). Runs that legitimately need longer should target a deployed/public --url instead of a local one.
If your account has too many tunnels open at once you'll get Tunnel limit reached — wait a minute and re-run; if the same local target is already being provisioned you'll get a short retry hint. Neither starts a run — just re-run once things settle.
Commands
| Command | What it does |
| --- | --- |
| seldon init | Scaffold a template seldon.yml |
| seldon cases add | Append a test case (--name --step … --criteria) |
| seldon cases | List the cases in the current seldon.yml |
| seldon validate | Validate the file locally |
| seldon run [--url <url>] | Ingest seldon.yml and run it; a local --url is auto-exposed via a Cloudflare tunnel (--auto-approve to skip the prompt); --no-wait to detach, --fail-on to gate, --label <tag> to profile |
| seldon report [--md \| --json] [--full] | Show the report for the last run (or --id <run>) |
| seldon logs --follow | Stream the agent's live reasoning |
| seldon runs | List QA runs with color-coded status |
| seldon cancel --id <run> | Stop a run mid-flight |
| seldon list | List baselines your key can see |
| seldon config · login | Manage credentials and settings |
| seldon api <operation> | Call any Blok API operation directly |
Run seldon --help or seldon <command> --help for the full reference.
Tagging runs for profiling
Attach freeform labels to a run — persistent ones in seldon.yml (labels: [...]) and per-invocation ones on the command line — to slice usage later:
seldon run --label nightly --label team-webLabels are merged (file + flags), de-duplicated, and shown in seldon runs alongside each run.
Staying up to date
seldon --version prints the installed version. When a newer release is available, the CLI shows a one-line upgrade nudge with the command that matches how you installed it — npm install -g @seldonqa/cli or brew upgrade seldon. The check is cached and never blocks a command; silence it with NO_UPDATE_NOTIFIER=1 (it's already quiet in CI and when output is piped).
Coding agents (MCP)
Seldon is also available as an MCP server, so coding agents can author baselines, run them, and read reports directly — the same report you see, rendered as Markdown — enabling a closed test-and-fix loop. See @blok/mcp.
Requirements
- Node.js 20 or newer
- A Seldon API key
Run seldon --help to get started.
