@unotest/web
v0.46.0
Published
AI-native E2E testing for web applications. MCP server (run_test / step / resume / inspect_runtime / agent_fix) + CLI runner + JavaScript DSL scenarios on a sandboxed AST engine + semantic DOM snapshots.
Readme
@unotest/web
AI-native E2E testing for web apps. You don't write the tests — your AI agent does, by driving your real app through the MCP server. You review and commit.
MCP server + CLI runner + JavaScript DSL + semantic DOM snapshots + structured failure bundles.
Full documentation — manuals, the box, CI, the judge — lives at
docs.unotest.com. This README is the short
tour; the package ships only guides/dsl-reference.md and
guides/agent-integration.md alongside it.
1. Setup
npx @unotest/web@latest initRun once in your project — any stack (Node, Django, Rails, Go, …), no
package.json required. init writes the config and a starter scenario,
sets up the browser (bundled Chromium, or your system Chrome if you pick it
during the prompt), and wires the MCP config so Claude Code / Cursor / Codex
pick the server up automatically. Re-run anytime to update — it never
overwrites your edits.
It also wires editor typings (unotest/jsconfig.json + generated .d.ts
files), so scenarios get autocomplete, hover docs and signature help in
VS Code / WebStorm instead of red squiggles. On an existing project this
self-installs on the first lint / e2e run after upgrading; refresh by
hand with npx @unotest/web types. Details: the "Editor setup" section of
node_modules/@unotest/web/guides/dsl-reference.md.
Testing a browser extension? launch in unotest.config starts the browser
with a profile (userDataDir) and extra arguments (args, e.g.
--load-extension=./ext; Chromium only). On a profile every scenario shares
one browser context — no isolation between them, and the run log says so.
profile: 'per-test' instead gives every test a new browser on a new,
empty profile: a suite of N files runs on N profiles.
2. Open the viewer
npx @unotest/web viewerA local IDE-style UI for your tests — no cloud, no account. This is your home base: browse scenarios, run them, and watch each step live. (Details in §5.)
3. Create a test — just ask your agent
In your AI editor (Claude Code, Codex, Cursor, …), open your project and describe the flow in plain English:
"Open https://playground.unotest.com, go to the Click vs Double-click section, click the Click me button, and check that a popup appears."
Through the MCP server the agent explores your live app — clicking,
filling, reading the real DOM via semantic snapshots — then writes a clean
scenario to unotest/e2e/<name>.js with stable selectors (getByTestId →
getByRole → … — never brittle CSS), runs it, and debugs itself if it
fails. You get a reviewable .js test: read it, tweak it, commit it.
For best results, point the agent at its authoring guide, shipped with the package at
node_modules/@unotest/web/guides/agent-integration.md.
4. Run a test
Click any scenario in the viewer, or run from the CLI:
npx @unotest/web e2e <name> # one scenario
npx @unotest/web e2e # list / usageRuns of one project take turns. A second run — from another terminal, the
viewer's Run button or an agent — joins a queue in unotest/.queue/ and
starts when the machine is free, instead of fighting the first one over
the same browser and the same seeded data. It costs nothing when nothing
else is running, the viewer shows who is waiting (and can cancel them),
and Ctrl-C gives up your place in line. Raise queue.concurrency in
unotest.config to allow more at once, or set UNOTEST_NO_QUEUE=1 for a
single command that must not wait. A collection is one run: its workers
still run in parallel inside it.
Running on a schedule
The easiest way: open the viewer, right-click a test or a collection →
Schedule… and pick when it runs. That saves to
unotest/schedules.yaml — a plain file, versioned next to the tests it
runs, so a box picks it up with your next push:
# unotest/schedules.yaml
schedules:
- collection: nightly
cron: "0 3 * * *"
- scenario: checkout/happy-path
cron: "0,30 4,16 * * *"
env: stagingHand-written setups can keep declaring schedules: [...] in
unotest.config — both sources merge, the yaml wins on a duplicate
target:
// unotest/unotest.config.mjs
export default {
schedules: [
{ collection: "smoke", cron: "0 * * * *" },
{ collection: "nightly", cron: "0 3 * * *", prepare: "node seed.mjs" },
],
};npx @unotest/web schedules prints the merged list, with a next …
preview per entry (your machine's clock) so a wrong cron is visible
before you push it. A hand-written expression outside the previewable
subset is marked not previewable — an executor may drop it;
npx @unotest/web schedules --check turns that mark into a non-zero
exit for a pre-push hook or a CI gate. Nothing in the CLI ticks: cron is
executed by whatever hosts your runs — your CI, or a box — so a clone of
your repo never starts running suites in the background. Run one now with
npx @unotest/web collection nightly --scheduled, or
npx @unotest/web e2e checkout/happy-path --scheduled for a per-test
entry (both also run the entry's prepare command, inside the same run
slot).
The viewer shows this same merged list everywhere — locally the yaml
half is editable and config entries are read-only ("declared in
unotest.config"); on a box the whole list is read-only: the box runs
whatever the deployed bundle declares, so the place to change a schedule
is your repo — edit, push, done. What you see locally is what a push
ships, 1:1. bundle push also records a fingerprint of this set in the
bundle, so if your config computes schedules from its environment, the
box notices the difference and says so instead of silently arming
another set.
Run on a box
Three commands, no token to copy:
npx @unotest/web login # a link and a code; approve in the browser
npx @unotest/web box create # rents a box named after the project, size Slogin signs this machine in to unotest cloud — the account that rents
boxes — and keeps the key in ~/.config/unotest/credentials.json
(mode 0600). box create creates the box with the environments the
suite already runs locally (each unotest/.env.<name> and its
APP_BASE_URL), waits for the box, writes
UNOTEST_BOX_URL to unotest/.env and the push and read tokens to
unotest/.secrets, pushes the suite and prints the viewer's address.
The tokens live thirty days; bundle push and box … re-mint them
through the cloud when a box says they have expired, and box access
does it on demand. box list shows every box on the account, box status
one of them, and box logs the end of a box's own log without logging in to
it; box destroy does what it says, and whoami and logout complete the
set. In CI, export UNOTEST_CLOUD_TOKEN instead of signing in.
Credit is added in the cabinet, by card, under "Add credit" on its home page
— never from the terminal. When the balance does not cover a new box's first
hour, box create says how much is missing (as far as the key may read the
balance) and where to add it, and exits 75 with nothing created and nothing
charged; run it again once the credit is there.
Anything that spends money or cannot be undone — creating a box, deleting
one, switching the Picker on — waits for a person unless the key is a CI
key. The CLI prints where to approve it and how long the request lives, waits
there, and then carries on exactly as if the call had gone straight through.
Nobody answers in ten minutes: exit code 75, nothing created and nothing
charged, run the command again. Somebody declines: exit code 2. --yes does
not change any of this — the flag says the terminal need not ask, not that
the account need not.
approvals lists what is waiting, and approvals <id> says what became of
one.
An agent does the same through the MCP server's cloud_* tools — listing,
renting, deleting, logs, the balance — with the same key and the same rules:
a rent or a delete comes back as a request to wait on, and
cloud_approval_wait waits for the person's answer.
A key carries permissions. login asks for enough to see the
account's boxes and work with them, never its money, and the approval
page in the browser is where you can narrow that further; login and
whoami print what the key ended up with. A key that may not read the
balance is still a working key — whoami says the balance is not
visible to it instead of failing — and a call it may not make comes
back as insufficient-scope, which no amount of signing in again will
change. Keys for an agent are issued in the cabinet, with the
permissions ticked there.
The Picker — grounding intent locators on the box's grounder — comes
with the box today: it lists at €0.05–€0.13 an hour depending on size,
and that is struck through, nothing is added to the bill. It is still
never switched on quietly, because that will not always be true.
box create --picker rents the box with it; without the flag a terminal
is asked once, with the price read from the cloud's own list — the CLI
has no price of its own — and the answer defaults to no. On a box you already have, box picker on and
box picker off switch it without re-creating anything: on writes
UNOTEST_GROUNDER_MODE=remote and UNOTEST_GROUNDER_REMOTE_URL to
unotest/.env and UNOTEST_GROUNDER_REMOTE_TOKEN to unotest/.secrets,
off clears the three again — unless they name a grounder that is not
this box's, which is left alone and reported. off also revokes the
box's Picker token, so a copy of it on another machine stops working
too: the grounder refuses it once it has re-read the operator's
revocation list, and off says how long that takes — "within about 2
minutes" — whenever the cloud knows both halves of the wait. Where it
does not, the sentence stays true and carries no figure rather than
half of one. Either way the charge follows at the
next hourly tick, because a started hour is charged at the rate it
started with. box status says which state a box is in, and — while the
Picker is on — which models the cloud grounds with: Our GPU or
External provider (Gemini). That is our grounder, not your
UNOTEST_GROUNDER_MODE; a cloud that has not heard from its grounder
prints nothing rather than a guess.
How a suite got installed is visible too. box envs prints, under
each environment, how long the box took to install its bundle and where
the packages came from — installed in 12.3s from this box's cache (252
of 252 from cache). box run <id> prints the same line for the bundle
that run came from, while the environment still serves it. The source is
the box's own measurement (what its npm cache already held of what the
lockfile names), not a reading of npm's output: on a fast registry a cold
install and a warm one take the same number of seconds, so the time alone
proves nothing. A box that says nothing about installs — one older than
this — prints nothing extra.
A box never updates itself. box status prints its version — what it
runs, and what its channel offers when the two differ. A new release is an
offer: the box keeps running what it is on until somebody presses Update
on its card in the cabinet, or an agent calls
POST /api/v1/boxes/<name>/update with the version it was offered (a
version that has since been overtaken is refused, with the one that is on
offer now). There is no box update command: the decision is a person's,
and the cabinet is where it is made. A cloud that reports no versions leaves
box status silent about them.
Without a box. An account can have the Picker on with no box rented
at all; the switch and the price for that live in the cabinet, and
picker on connects THIS project to it — it fetches the project's own
access and writes the same three keys, keeping the token out of the
repository. It says the account's monthly ceiling before anything is
spent: there is no hourly charge for a Picker without a box, so the
ceiling is the price. The figure comes from the service, which takes it
from the grounder that enforces it — and when the grounder has not
reported one, the command says that instead of guessing. picker off disconnects this project and clears the three
keys — leaving the Picker on for the account, because the switch
that turns it off for everything, and stops it counting against your
monthly ceiling, is the cabinet's. A box's own Picker stays
box picker.
An agent can drive all of it without a terminal: no prompt is shown without a TTY, and the exit codes say what happened — 75 the payment was not confirmed in time, 76 the terms of service need accepting (a person must, in a browser), 77 not signed in.
Sending the suite to a box
If your team runs a box (a machine that keeps the environments, the schedules and the history in one place), tests get there by being pushed:
npx @unotest/web bundle push --box https://tests.example.comIt packs unotest/ — the suite package, with its config, its
package.json and its lockfile — and nothing else. Your application's
manifest and dependencies do not travel and are never installed on the
box, which is why a box can run a suite for a project whose own
dependency graph it could not resolve in the first place.
The box never reads your repository: it holds no key to it, so
nothing is there that you did not send. Uncommitted work is included and
marked wip, which is what makes this a fast loop rather than a release
step. Secrets stay behind: unotest/.env* and .secrets* never travel,
because environment values belong to the environment and are injected
over the bundle when it runs. Install hooks (postinstall, prepare,
...) are removed from the packed unotest/package.json: a box installs
the suite's dependencies, it does not run a checkout's hooks. Your
DEPENDENCIES' install scripts do run — a native module builds or fetches
its binary as usual — because on a box that install happens in a
container of its own, with the bundle tree as its only writable mount
and none of the box's state in reach.
What would break on the box is refused here instead — a scenario that
reads a file outside unotest/, a missing unotest/package-lock.json,
no @unotest/web pin. The bundle's id is a hash of its content, so
pushing an unchanged suite is a no-op.
From a CI job, push and run in one command:
npx @unotest/web bundle push --run --env test --collection smoke --pr 412The box answers as soon as the runs are queued and gives you their ids —
a suite takes minutes, and a request held open that long is a timeout,
not a result. Watch them in the viewer, or let the box report the checks
back to GitHub. --pr makes a newer push withdraw the older runs of the
same pull request that are still waiting, so three pushes in five minutes
cost one suite. Recipe: Pushing suites → A minimal CI job.
The values a suite runs with on the box — APP_BASE_URL, its .env
settings, its secrets — are sent separately, from the same files a local
run reads:
npx @unotest/web env push dev # .env + .env.dev + .secrets + .secrets.dev
npx @unotest/web env set dev REGION eu-west
printf '%s' "$KEY" | npx @unotest/web env set dev API_KEY --secret.env* become the environment's variables, .secrets* its secrets,
APP_BASE_URL its target; UNOTEST_* stay home. A secret's value never
prints; env push --dry-run shows what the other values resolved to. The
box and the token come from the same places bundle push reads them
from — flags, the environment, then unotest/.env and
unotest/.secrets — and the token must be minted with --values on the
box. Admins also see and change them on the box's admin page. Manual: An
environment's values on a
box.
A run on the box happens inside the box's own container: nothing on your
machine is reachable from it — no localhost, no kubectl port-forward,
no tools installed on your laptop. Preconditions must probe the
environment's public URL, and assertJudge runs its judge in-process there
(UNOTEST_JUDGE_MODE=local, @unotest/judge in the suite's
package.json, the API key as a secret). Details:
Judge on a box and
What a box cannot reach.
Reading results from a box
bundle push --run gives you a run id. Read what became of it from the
same terminal — no browser, no shell on the box:
npx @unotest/web box runs --env acme/test --latest
npx @unotest/web box run checkout-mqf3pwr1 --env acme/testbox run prints the failure, the results of soft steps and the judge's
verdicts. When you want the artifacts too:
npx @unotest/web box run checkout-mqf3pwr1 --env acme/test --downloadThat saves the run's *.unotest.zip under .unotest/box/ and unpacks its
failure bundle into .unotest/failures/ — where list_failures,
get_failure_* and agent_fix already look, so your agent debugs a box
failure with the commands it uses for a local one. Add --no-screenshots
to leave the step frames on the box; they are usually most of the bytes.
Reading uses a personal read token, not the project one. Put it in
unotest/.secrets, next to your other project secrets:
UNOTEST_BOX_READ_TOKEN=ubr_...and the box's address in unotest/.env as UNOTEST_BOX_URL. Both files
are read directly, so this works the same from the terminal and from your
agent. Exporting either value in your shell works too and wins over the
file; --box and --token win over both.
The token is never read from unotest/.env — that file travels inside a
pushed bundle, and .secrets does not.
unotest/.secrets is created owner-only (0600), so the tokens in it are
not readable by other users of a shared machine. A file left wider by an
earlier version is narrowed the next time a command writes to it, and the
command says so once. Editing a secret in the viewer's variables panel
keeps the same width. unotest/.env is configuration, not secrets, and
keeps whatever width your project gave it.
On Windows there are no POSIX permission bits, so nothing is narrowed there and no width is promised: protect the file with the tools your system has.
Keeping it in the file is what lets your agent read runs: an MCP server is started by your editor, so a token that only exists in your shell is a token the agent never sees. A token added while the server is running is picked up on the next call — nothing to restart.
With a box rented through box create, the token is minted for you —
box access re-mints it — and lives thirty days. On a box set up by
hand you mint it for yourself on the box's guard, where it is shown once
and can be revoked. Either way it is read-only whatever your role on the
box — it cannot start a run or change a value. UNOTEST_BOX_TOKEN is the
project's push credential and cannot read runs. box envs lists the
environments the box serves — an empty list means the box has none yet,
not that the token is too weak to see them. The commands that name one
(box runs, box queue, box run, …) say the same thing rather than
sending the name to a box that has nothing to match it against: a box
that names none has none configured, which is ours to put right. A run
you have just ordered may take a moment to appear.
5. Watch it run — the viewer
npx @unotest/web viewer opens a local UI (no cloud, no account):
- Scenario tree — scenarios, collections, helpers, and run history.
- Run from the UI — results stream live over WebSocket as steps execute.
- Block view — each
step(...)group folds into a section with per-step status; a failure pins an error card to the offending line. - Step debugger — set gutter breakpoints, run paused, inspect the live DOM at each stop.
- Inspector — screenshots and failure artifacts per run.
- Integrated terminal — a docked xterm (toggle
⌃`) to drive the CLI without leaving the window.
Writing by hand (optional)
Scenarios are plain .js. Every executable step sits inside a
step("what this does", () => { … }) block — plain-English intent the agent
reads back to repair a step when a selector drifts:
function test_login() {
step("Log in as the demo user", () => {
goto('https://app.example.com');
fill(getByLabel('Email'), '[email protected]');
click(getByRole('button', { name: 'Continue' }));
});
step("Dashboard is shown after login", () => {
assertText(getByRole('heading'), 'Dashboard');
});
}The DSL uses the familiar modern browser-automation vocabulary, so it reads the way you already expect:
- Navigation —
goto,reload,waitForUrl - Locators —
getByRole,getByTestId,getByLabel,getByText,getByPlaceholder,locator - Actions —
click,fill,press,hover,check,uncheck,selectOption - Assertions —
assertText,assertUrl,assertVisible,assertCount,assertValue,assertJudge(LLM-judge for free-form text, backed by@unotest/judge) - Setup —
dbQuery,dbExec,apiCall,shell,evaluate
Full reference, shipped with the package:
node_modules/@unotest/web/guides/dsl-reference.md (plus operational guides
alongside it in guides/).
Ecosystem
| Package | Role |
|---|---|
| @unotest/web | this package — CLI / MCP server / runner |
| @unotest/viewer | local results browser |
| @unotest/protocol | shared types |
| @unotest/dsl | scenario parser + validator engine |
| @unotest/eslint-plugin | the same scenario diagnostics through ESLint |
License
MIT — see LICENSE.
