@pwtap/mobile-inspector
v2.3.0
Published
Mobile test recorder for the Playwright Test Automation Platform — local service, browser UI and code generator. A dev tool: the runtime contracts live in @pwtap/mobile-core
Readme
@pwtap/mobile-inspector
Mobile test recorder for the Playwright Test Automation Platform. Tap around a real device, get a Playwright test you can run and commit.
It is a dev tool: nothing here is loaded by a test at runtime. The contracts a generated test needs live in @pwtap/mobile-core, and the device work is done by whichever driver plugin you have installed — @pwtap/plugin-maestro or @pwtap/plugin-appium.
Open it
npx mobile-inspect # the current project
npx mobile-inspect ./my-app # or a specific oneIt prints a loopback URL and opens it in an app-mode window using the Chromium Playwright already installed — there is no Electron in the dependency graph. If no browser is available it just prints the URL, which is an equally usable inspector in any browser. One inspector per project: a second launch is refused and points at the first.
Record a test
- Connection → pick a driver, platform, device and app (installed apps are listed for the selected device; you can also point at a local
.apk/.app/.ipa/.zipor anhttps:artifact URL). - Tap the device screen to drive it; ⌘/Ctrl + tap to record the step. Dragging works the same way — a plain drag swipes, a modifier-drag records the swipe. Getting to the screen you came to record takes the same clicks as recording it, so the two are separated: turn on Record in the Device panel if you want every interaction written down (the modifier then does the reverse for one gesture). Back / Home / Enter / Screenshot sit under the viewport, under the same rule.
- Right-click any element for its attributes and its ranked locator alternatives, with Tap / Fill / Long press / Wait / Assert visible / Assert not visible / Is visible / Screenshot / AI assert, scrolling inside that element, copy-as-code and reveal-in-tree. Anything the connected driver cannot do is disabled with the reason, and a checkbox writes a step down without running it.
- The editor is yours. Type in it and the recording splices new actions into what you wrote instead of overwriting it; completion offers
mobileApp's methods and locators for the elements on screen right now. Undo/redo work on the timeline, not just the text — and clicking a recorded step shows the screen it produced. - Run executes the draft through the project's own Playwright, in the driver's project with its gate variable set, and streams the output back.
- Save… writes a new file (browse the project's real directories) or appends to an existing recording, merging imports and keeping the file's existing tests.
A recording is saved with its driver's extension — *.maestro.ts or *.appium.ts — because the extension is what decides which Playwright project collects the test, which env var gates it and which timeout applies. Saved into the wrong one, a test would silently never run; git mv to the other extension is all it takes to move it.
What a generated test looks like
import { test, expect } from '@fixtures';
test.use({
mobileTarget: {
driver: 'maestro',
platform: 'android',
device: 'Pixel_7_API_34',
appId: 'com.example.app',
},
});
test('sign in', async ({ mobileApp }) => {
await mobileApp.tap({ accessibilityId: 'emailField' });
await mobileApp.fill({ accessibilityId: 'emailField' }, '[email protected]');
await mobileApp.tap({ text: 'Log in' });
await expect(mobileApp).toBeVisible({ text: 'Dashboard' });
});Driver-neutral by construction: the same test body runs under either driver — mobileTarget.driver is the only thing that changes. The device is pinned by its stable name (an AVD name, a simulator name), never the ephemeral adb serial, so the test still finds the device after a reboot.
Locators
Every element is offered as a ranked list rather than one guess, scored by how durable it is: accessibility id → resource id → exact text → index-qualified text → coordinates, with penalties for non-unique matches and for elements that belong to another app (a status bar, a system dialog). The confidence badge and the warnings are the point — a coordinate locator works today and breaks on the next layout change, and the UI says so instead of hiding it.
MCP server
The same package serves an MCP server, so an agent can drive a device through the contracts your tests already use:
npx mobile-mcp . # or: npm run mcp:mobile in a project with a mobile pluginNine tools — mobile_drivers, mobile_devices, mobile_connect, mobile_disconnect,
mobile_hierarchy, mobile_locators, mobile_screen, mobile_perform, mobile_codegen. The one that
justifies the server is mobile_locators: it returns the ranked, uniqueness-checked, fragility-annotated
list described below, which no shell command produces. adb shell uiautomator dump gives raw XML with no
scoring, and an agent working from that writes coordinate taps.
Add it to any MCP client with the block npx @pwtap/create mcp prints, or install the pwtap Claude Code
plugin, which derives it from the plugins you have.
Acting on the device is off by default. mobile_perform stays listed and refuses, naming the switch
(PWTAP_MCP_ALLOW_ACTIONS, or the plugin setting) — hiding the tool only teaches a model to reach for
adb shell input tap through your shell instead. There is deliberately no shell, adb, simctl,
uninstall or erase tool at all: an MCP tool is approved by name, once, so one of those allowed is a
permanent unaudited escape from the permission gate that does see the real command string.
Screen text, ids and labels are quoted to the model inside a per-call <device-material-…> tag, because a
login form is attacker-controlled input. mobile_screen returns a file path by default — a screenshot
of a logged-in app is a credential — and the hierarchy is depth- and count-bounded.
At most one device session per process, and the server never takes the device lock itself: the adapters do,
inside connect()/close(). An idle session closes on its own after PWTAP_MCP_IDLE_MS (default ten
minutes), so a forgotten agent session cannot block your own test run. Design notes and the decision log:
docs/mcp-plan.md.
Trust boundary
The service binds to loopback on a random port with a per-launch token, and treats the browser as untrusted even though it is local: every command is validated field by field, file writes and directory listings are confined to the project (by path segment, and following symlinks), the artifact path is validated before it reaches an installer, and runs are argv-only spawn with an explicit environment.
The token stays out of anything that keeps it. The window carries it in an x-inspector-token header, so the address printed on launch has no credential in it, and the single-instance lock file holds only a port and a pid. The one place a token is still printed is when no browser could be opened and you have to paste the URL into your own — that line says so. Closing the window releases the device lock, kills any run and removes the temp files; merely losing the connection does not — reload the page and the recording is still there.
Design
The decision record is docs/mobile-inspector/architecture.md in the repository: 14 ADRs covering the loopback-service host, the driver-neutral action IR, node identity, the capture schedule, the trust boundary and the dependency policy.
License
MIT
