@deeeed/metamask-harness
v0.69.0
Published
One CLI for operating MetaMask Extension, Mobile, and Core and producing reviewable recipe evidence. Run it inside a checkout; the product, slot, ports, and runtime paths are detected automatically.
Readme
mm-harness
One CLI for operating MetaMask Extension, Mobile, and Core and producing reviewable recipe evidence. Run it inside a checkout; the product, slot, ports, and runtime paths are detected automatically.
- Action: one typed operation.
- Recipe: a reusable, parameterized graph of actions and called recipes.
The generic graph engine and evidence schemas live in Farmslot packages.
mm-harness owns MetaMask runtime control and domain capabilities.
Daily and release QA
The same recipes can validate yesterday's domain changes together or a release
that is behind main. The mms-recipe-qa skill freezes the change scope and
criteria; the team library supplies smoke and reusable journeys; this CLI runs
them on the selected app and retains assertions, traces and captures.
Use the installed skill's references/domain-qa.md for window/range inputs and
report policy. Scope selection and scheduling are outside the harness. Discover
recipes through run --list and run <recipe> --describe, then follow the
locked version's help --json for composition and drift diagnosis. Adapt a
recipe to the tested release without changing the release or weakening its
criteria. An internal action that needs development instrumentation cannot
prove behavior on an unmodified release artifact.
Getting started from zero
No checkouts yet? mm-harness setup-base clones the MetaMask product repos into one
standard layout and runs each repo's own dependency install — no install of this
package required to start:
npx -p @deeeed/metamask-harness mm-harness setup-baseIt produces, under ~/dev/metamask by default:
metamask-extension-1..N metamask-mobile-1..N core-1..NRun without flags in a terminal, it asks which projects you want and how many copies of each. Pass flags to skip the prompt entirely — which is how scripts, CI, and parent tools should invoke it:
mm-harness setup-base --only core --counts core=1
mm-harness setup-base --dir ~/work/metamask --counts extension=3,mobile=1,core=2
mm-harness setup-base --only mobile --dry-run # plan only, changes nothingThe base directory resolves by precedence: --dir, then $MM_HARNESS_BASE_DIR,
then saved preferences, then ~/dev/metamask. A successful run saves the base
directory and counts it used, so later runs adopt them — inspect with
--show-config, clear with --reset-config. Re-running is safe: existing
clones are fetched, never reset and never deleted.
Scope fence. It clones and installs dependencies. Nothing else — no platform
toolchains, no simulators, no .env files, no builds. Workflow onboarding is a
prompt; environment readiness is mm-harness doctor, which is exactly where it
points you next.
setup-base is a normal command on the single mm-harness entrypoint. Its work
stays in one readable shell leaf, and npx -p makes that same command available
before a global install.
Install
npm install -g @deeeed/metamask-harness@latest
mm-harness --version
mm-harness doctordoctor is read-only. doctor --fix repairs harness-owned runtime state but
does not launch an app, invent credentials, or choose a wallet fixture.
mm-harness doctor --fix
mm-harness fixtures init --from /secure/path/wallet-fixture.json
# Disposable public testing only; never fund this wallet:
mm-harness fixtures init --devUse mm-harness update to update a published installation.
Operate
# Extension
mm-harness launch
mm-harness launch --sidepanel
# Mobile
mm-harness launch ios
mm-harness launch android
# Any checkout
mm-harness status
mm-harness stop
mm-harness logs
mm-harness debug
mm-harness reload # Mobile Metro reload / Extension CDP refresh
mm-harness fixtures set
# Delete the current wallet, then reapply the canonical fixture from clean state.
mm-harness fixtures resetExtension launch keeps its incremental watcher running. Refresh the active
page after a successful rebuild; use launch --build only for an explicit
production-like LavaMoat rebuild.
Remote feature flags
--remote-flag KEY=VALUE[,KEY=VALUE] pins a remote feature flag for the launch,
which is how a screen gated behind an A/B variant becomes reachable:
mm-harness launch --remote-flag perpsTAT3938AbtestScreenVsBottomSheet=treatment
mm-harness launch ios --remote-flag perpsEnabled=true,perpsConfig='{"tier":2}'A bare word is a named variant (treatment becomes { "name": "treatment" }),
true/false is a boolean, and a value opening with { or [ is parsed as
JSON. The grammar is parsed once and every leaf receives the resolved map, so a
boolean stays a boolean rather than becoming the string "true".
Every launch makes the app's override set exactly the wallet fixture's
remoteFeatureFlags object (the baseline every run shares) merged with the
--remote-flag map (a one-launch override), on both clients. A plain launch
therefore drops a previous run's --remote-flag variant and keeps the fixture's.
Mobile clears and re-applies the set through the running
RemoteFeatureFlagController and fails the launch if it does not read back. The
Extension writes it into the per-run dist snapshot manifest and re-reads the
harness-owned record there (_harness.remoteFeatureFlags, written on every
snapshot), so the launch reports the pin it proved; reuse of a running browser
is allowed only while that record already equals the request, a reload in
place carries the record across the rebuilt snapshot, and --watch (which
never re-creates the snapshot) refuses --remote-flag. A recipe's own recovery
relaunch carries the live pin rather than dropping it mid-run. A verified
release artifact is never rewritten, so a launch that would have to pin it
(--remote-flag or a fixture that declares flags) refuses instead of testing
the wrong arm.
Read or change flags on a running app with the actions. fixtures set follows
the same rule as launch: the fixture's block is the baseline and
fixtures set --remote-flag KEY=VALUE layers a one-run override on it, which is
how prepare --remote-flag keeps its pin through the fixtures step.
mm-harness call metamask.feature_flags.read
mm-harness call metamask.feature_flags.set --arg 'flags={"perpsEnabled":true}' # mobile
mm-harness call metamask.feature_flags.clear # mobiledoctor and status print a feature flags: line naming every active
override, so a slot never runs a proof under an unnamed variant. The line
distinguishes (no runtime) and an unreadable snapshot from a genuine
0 override(s).
Log sources stay separate:
mm-harness logs --source extension # Extension page
mm-harness logs --source dapp # active dapp
mm-harness logs --source webpack # compiler
mm-harness logs --source watcher # watcher lifecycle
mm-harness logs --source rebuild # incremental rebuilds
mm-harness logs --source app # Mobile app
mm-harness logs --source metro # Mobile bundlerLive Mobile and Extension recipes automatically index a bounded, redacted
network/run-summary.json with request timing and recipe-node boundaries. Use
explicit app.network_capture and app.network_assert nodes for filtered
windows and machine-checked request expectations. See
Recipe-scoped network capture.
Mobile and Extension recipes can bracket explicit UI-smoothness windows with
app.performance_capture and verify their evidence with
app.performance_assert. Capture is never automatic. See
UI smoothness capture.
Work a task
# Pick a checklist from the shared catalog.
mm-harness execution-template list --package-templates <catalog> --json
# Write the task directory the agent works from.
mm-harness task init temp/tasks/fix/TAT-1 \
--flow fix-bug --template fix-bug/mobile \
--package-templates <catalog> --title "Wrong total on the receipt" --ticket TAT-1
# Bring the checkout to a proven-ready state before the first step.
mm-harness prepare --artifacts-dir temp/tasks/fix/TAT-1/artifacts --json
# Mark progress; the task's own ./mark shim routes here.
mm-harness checklist mark temp/tasks/fix/TAT-1 start
mm-harness checklist mark temp/tasks/fix/TAT-1 1
mm-harness checklist mark temp/tasks/fix/TAT-1 complete --mark-lasttask init writes TASK.md, CHECKLIST.md, the mark shim, and inputs/.
CHECKLIST.md is the only file whose boxes count as steps. It writes no
checklist-target.json; absent means CHECKLIST.md + SIGNAL.json.
prepare runs doctor --fix, status, launch --verify, fixtures set, and
verify in order, stops at the first failure, and records one readiness record
at <artifacts-dir>/sandbox.json. Core skips launch and fixtures; Mobile takes
ios/android from the checkout's configured target, then from the devices
status reports, and asks when that is ambiguous. Exit 0 means
ready. See the task directory.
Discover and prove
mm-harness actions positions
mm-harness actions --action metamask.wallet.ensure_unlocked
mm-harness call metamask.wallet.ensure_unlocked
mm-harness run --list
mm-harness run perps.smoke --describe
mm-harness run wallet.smoke --describe
mm-harness run wallet.import method=auto
# When a visible Terms of Use gate follows import and acceptance is authorized:
mm-harness run wallet.accept-terms accept=true
# Use visible client UI to delete and import the wallet again.
mm-harness run wallet.reset-import
mm-harness run path/to/recipe.json market=ETH --plan
mm-harness run path/to/recipe.jsonrun selects a checkout-local artifact directory unless
--artifacts-dir <dir> overrides it. Human output prints diagnostics and
absolute paths to the report, trace, executed recipe, and artifact manifest.
Use mm-harness last --json to resume without repeating the last operation.
For automation, --json emits one stable document. run --json-stream emits
line-buffered JSONL progress and a terminal event. Its status matches the recipe's
pass, fail, or unknown verdict; unknown still exits nonzero.
Team libraries
export RECIPE_LIBRARY_PATH="team=$HOME/shared-library/team-recipes"
mm-harness run --listShared libraries hold durable actions and composable recipes. Task acceptance criteria remain task-local. See Recipes.
For optional model advice matching acceptance criteria to library flows and assessing
drafts or evidence, see Optional recipe advisor.
Select a provider with --advisor typesafe or save advisor.enabled and
advisor.provider through mm-harness config. Inspect the effective policy with
mm-harness recipe-quality advisor-status --json; use --advisor off per task.
Ordinary recipe execution needs no advisor or API key.
Exporting is not required: a library registered with mm-harness config set
libraries.<name> <path>, or checked out beside the product checkout, reaches
run, actions, and template discovery too. mm-harness status lists what this
machine resolves, with each location's revision and recipe count.
Teams maintain their library checkouts. Update them before starting a task; the harness records local revisions and recipe digests but does not fetch or check for upstream updates. Keep the selected library unchanged during a run.
Publish a change
# Write artifacts/pr-description.md in the repository PR template shape (prose
# only), then render the publishable body.
mm-harness pr-body render temp/tasks/dev/TAT-1 --command 'mm-harness run artifacts/recipe.json'pr-body render writes artifacts/pr-body.md: the description with the "Validation
Recipe" and "Validation Logs" sections rendered from artifacts/recipe.json and
artifacts/recipe-run/report.md (and "Recipe Workflow" from artifacts/workflow.mmd
when it exists), replacing any pasted by hand. The worker never pastes artifacts
into markdown, so fences cannot break.
Review a change
# Which team library owns the change (declared --domain wins, else owned-paths.json).
mm-harness domain
# The review guide: base review + the library's anti-pattern families and parity rule.
mm-harness help review --domain perps
# A checklist the mark shim can track: Setup, Base review, Domain patterns, Parity, Verdict.
mm-harness review checklist --domain perps --out temp/tasks/review/TAT-1/CHECKLIST.md
mm-harness review checklist --domain perps --since <last-reviewed-sha> # incremental re-review
# Where this machine keeps libraries and the other platform's checkout (outside any repo).
mm-harness config set references.extension ~/dev/metamask-extensionA team library supplies review/antipatterns.md (one ## section per family),
optional review/antipatterns.<adapter>.md families appended for that adapter only,
review/parity.md, and owned-paths.json; the harness composes them and never
copies them. Locations resolve env (RECIPE_LIBRARY_PATH, MM_HARNESS_REF_<ADAPTER>)
→ mm-harness config → sibling directory beside the checkout → shallow fetch into
<checkout>/.skills-cache/<name>. A missing reference checkout degrades the parity
step to "not checked" with the reason.
Recover
mm-harness doctor
mm-harness doctor --fix
mm-harness verify
mm-harness cleanupFailures name one exact next action. Core is headless; browser, device, logs,
and debugger capabilities are reported as unavailable instead of fabricated.
Stable exit codes are 1 runtime/action failure, 2 invalid CLI usage, 3
infrastructure failure, 4 bounded recovery refusal, and 5 validation/trust
failure.
Reference
- Recipes — discover, compose, author, and share proof.
- Security — trust, approval, fixtures, and evidence safety.
- QA — clean-machine and human release checks.
- Contributing — ownership, layout, and change gates.
For development, point the installed command at a source checkout:
export MM_HARNESS_BIN=/path/to/metamask-harness/bin/mm-harness
mm-harness --versionUnset MM_HARNESS_BIN to return to the published installation.
Workflow entry points and completion
Explicitly invoke mms-recipe-dev, mms-recipe-fix-bug, mms-recipe-qa or the team's static review skill. mms-recipe-cook selects or resumes their shared Markdown checklist. Matching descriptions or changed paths never activates these workflows.
Static Perps review uses a generated source checklist without installing or launching the app harness. recipe-qa freezes a PR, range or daily change window and requires runtime smoke plus selected recipe proof. Official release validation uses its named release profile and proof lane.
TASK.md holds the request; CHECKLIST.md is the sole step list. The frozen terminal contract and task-local mark validate completion. Publication, retained sessions and workspace cleanup belong to the caller. Farmslot handles them through its control plane; direct invocation returns the outcome to the engineer. See the task-directory guide and QA tutorial.
Portable runtime proof
validateRuntimeProof(bundle, smokeTarget) validates a retained Recipe v1 artifact package, requiring executed behavioral assertions for its declared targets. A passing command or end node alone is insufficient. QA skills use this exported function through their locked harness installation; no gateway, farm profile or run ID is required.
