@minns/browser
v0.5.2
Published
A browser an agent can use by saying what it wants, and repeat reliably: the page as an outline, steps that find their element again before acting, submits that wait for approval, data checked against a JSON Schema across pages, and routines and extractio
Maintainers
Readme
@minns/browser
A browser an agent can use by saying what it wants, and a way to do the same task again, reliably.
import { chromium } from "playwright-core";
import { BrowserPilot, PageDriver } from "@minns/browser";
const page = await (await chromium.launch()).newPage();
const pilot = new BrowserPilot({ driver: new PageDriver(page), model });
pilot.startRecording();
await pilot.goto("https://shop.example/login");
await pilot.act("fill %email% into the email field", { variables: { email: "[email protected]" } });
await pilot.act("fill %pw% into the password field", { variables: { pw: { value: secret, secret: true } } });
await pilot.act("click Continue");
const total = await pilot.extract("the order total", { schema: { type: "object", required: ["total"], properties: { total: { type: "string" } } } });
const routine = pilot.stopRecording({ name: "sign in and read the total" });
// Later, for anyone: no model calls while the pages still match.
const run = await pilot.replay(routine, { email: "[email protected]", pw: { value: bobs, secret: true } });model is anything with complete(messages) => Promise<string> (see
src/model.ts); @minns/agent-forge providers fit as they are. A cheaper
model can take the smaller questions: models: { extract, heal } (also
act, observe) gives one per purpose, and every call still says its
purpose in metadata.purpose (browser.act, browser.observe,
browser.extract, browser.heal). A purpose with no model of its own is
asked of model; heal (a replay finding a changed step's element again)
falls back to act first.
Contents
Architecture
Two halves. BrowserPilot thinks and holds the model; a BrowserDriver
holds the page and never calls a model. PageDriver is the driver for a
Playwright page in this process, and the two can live apart: the minns
workspace box runs the driver, the agent runs the pilot, and model keys stay
out of the browser's reach.
flowchart TB
subgraph Agent["Agent (any model)"]
A[Instructions in words<br/>act, observe, extract, routines]
end
subgraph Pilot["BrowserPilot (thinks; holds the model)"]
P1[prompts.ts<br/>page inside page_content]
P2[routine.ts<br/>record, replay, heal, versions]
P3[extract.ts, recipe.ts<br/>schema, pages, recipes]
P4[network/learn.ts, call.ts<br/>API learned and validated]
P5[memory hooks<br/>hints, learn]
end
subgraph Driver["BrowserDriver / PageDriver (holds the page; no model)"]
D1[snapshot.ts<br/>outline, ids, targets]
D2[resolve.ts, fingerprint.ts<br/>the same element again]
D3[pointer.ts, typing.ts<br/>a hand's moves and keys]
D4[popups.ts, notices.ts<br/>what stands in the way]
D5[forms/<br/>ledgers, fills]
D6[apps/<br/>grids, tables, documents]
D7[network/<br/>recorder, transport, caches]
D8[guards.ts<br/>URL policy]
end
subgraph World["world.ts: our isolated world per frame (no footprint)"]
W[Page.createIsolatedWorld<br/>one DevTools session per target]
end
subgraph Browser["Chromium (Playwright, CDP)"]
B1[Pages, frames, shadow roots]
B2[Network service<br/>TLS, HTTP/2, cookies]
B3[Cache Storage, IndexedDB]
end
Site[(The site)]
Model[(Model)]
Store[(Routines, recipes,<br/>captures: plain JSON)]
A -->|say what you want| Pilot
Pilot <-->|look / perform / inspect<br/>form / fillForm / callApi / readUrl| Driver
Pilot <-->|complete messages| Model
Pilot <--> Store
Driver --> World
World --> Browser
Driver -->|Input.dispatchMouseEvent<br/>Input.dispatchKeyEvent| Browser
Browser <-->|the browser's own TLS and HTTP/2| Site
D7 -.->|fetch from our world of the tab| B2
D7 -.->|CacheStorage, IndexedDB domains| B3Three rules shape every box:
- The page is data. Everything read off a page reaches the model inside
<page_content>, which the page cannot close; what the driver says about a page quotes its words rather than speaking them. - The driver leaves no footprint. Nothing kept between calls lives in the
page's world (
world.tsholds it in an isolated context per frame), no attribute is written to the DOM, no window name added, no prototype or DOM method changed.tests/footprint.test.tschecks it. - Nothing is guessed. A step never presses a control it cannot identify with confidence; a value the page does not show is null; a recipe or an API is kept only while a direct read gives back what the model read.
How an agent uses it
An agent says what it wants in words. The pilot shows the model the page as an outline, gets a step back, and the driver finds the element again and acts on it like a hand. Everything the agent does once can be recorded as a routine and replayed without a model while the pages still match.
sequenceDiagram
autonumber
participant Agent
participant Pilot as BrowserPilot
participant Model
participant Driver as PageDriver
participant Page as Chromium / site
Agent->>Pilot: act("click Continue", { variables })
Pilot->>Driver: look({ forAgent: true })
Driver->>Page: accessibility tree + DOM + layout (one call per frame)
Driver-->>Pilot: outline with ids, notices, changes since last look
Note over Pilot: a bot check or a refusal stops here: reason "blocked"
Pilot->>Model: prompt with the outline inside page_content
Model-->>Pilot: { step: { method: "click", id: "0-53" } }
Pilot->>Driver: perform(step)
Driver->>Page: popup watcher: anything new over the page?
Driver->>Page: find the element again (fingerprint), wheel it into view
Note over Driver: a submit without allowSubmit returns reason "submits" with expect
Driver->>Page: Input.dispatchMouseEvent (a hand's path, rest, hold)
Driver->>Page: wait for the fetches the click started, then a quiet DOM
Driver-->>Pilot: StepOutcome (target, found, closed banners, status)
Pilot-->>Agent: outcome (and the step, when recording)The same loop runs observe (a question about the page) and extract (data
by schema, every part of the page, over pages). A replay runs the recorded
steps through the driver alone: the model is asked only to find a renamed
element, and a step that lands elsewhere, meets a sign-in page or fails a
verify stops the replay with the cause and no model call.
sequenceDiagram
autonumber
participant Agent
participant Pilot as BrowserPilot
participant Driver as PageDriver
participant Site
Note over Agent,Site: Learning a site's API once
Agent->>Pilot: startCapture(); goto(list page); extract(schema)
Pilot->>Driver: record the network (requests, who sent each, scripts, streams)
Agent->>Pilot: stopCapture(); learnApi({ target })
Pilot->>Pilot: rank responses, feeds, platform views, cached answers
Pilot->>Pilot: trace every value: page, storage, cookie, response, bundle, IndexedDB
Pilot->>Driver: callApi(candidate) from our world of the tab
Driver->>Site: the browser's own fetch (its TLS, HTTP/2, cookies, referer)
Site-->>Driver: rows
Pilot-->>Agent: recipe, kept only because the rows match the browser's
Note over Agent,Site: Every replay after
Agent->>Pilot: replay(routine with an api step, { validators })
Pilot->>Driver: open the learned page if the tab is elsewhere; callApi(recipe)
Driver->>Site: If-None-Match ... 304, or the rows
Pilot-->>Agent: data with no browser steps and no modelWhat it does
Seeing the page
The page as an outline. The browser's accessibility tree joined to the DOM: roles and names as a person reads them, one element per line with an id, iframes (cross-site ones too) and shadow roots included. Text has no id (it is read, not pressed) and a short one sits on its element's line. It says what things are, not how they are drawn.
[0-2] RootWebArea: Checkout
[0-24] heading: Checkout
[0-9] paragraph: Total £84.20
form
[0-5] textbox: Email value="[email protected]"
[0-6] select value="UK"
[0-37] option: UK [selected]
[0-53] button: Pay now
[0-64] iframe
[4172-5] button: Card button
(40 more lines below the view: 32 links, heading "Help"; scroll down, or expand 0-91 to read them)Viewport-first, and only what changed. A look writes what is on screen
and just around it in full, and sums up the rest a section at a time, which
a scroll or look({ expand }) brings in; look({ full: true }) reads every
line. A look of the same page as the reader's last one also carries
changes: the lines added, removed and changed, under a header that counts
them, when that is well under the whole. look({ forAgent: true }) works the
changes out against what the agent was last shown in that tab, never against
a look the pilot took for itself. On a shop listing that is about 2,200
tokens a look instead of 7,100, and a checkout filled in over seven steps
costs a fifth of what whole outlines did (see CHANGELOG.md).
Every step finds its element again. Before acting, the driver reads the
page afresh and finds the element by where it was and what it was (a
fingerprint: tag, role, name, test id, a short path). A different control that
took the recorded one's place is refused, not clicked. A positional step
takes the element exactly there of the same tag and role, which is how a
loop over a list's rows takes each row's control.
The document's status and headers. Each tab keeps the HTTP status and headers its document came with, so a 403 or 429, and what a guarding vendor said in its headers, are known before the page has drawn a word.
Acting like a hand
A pointer with a hand's path. Behind the pointer option (on by default;
pointer: false keeps Playwright's own actions), the mouse moves along a
path with Fitts's law timing, an early velocity peak and a long tail, a bow,
an overshoot and correction past 200 px, a slow tremor, in whole pixels, and
rests 100 to 300 ms before a press held 40 ms or more. It keeps its place per
page between steps. A target below the fold is rolled into view a wheel
notch at a time (about 100 px with gaps), and a scroll step rolls the page or
a box the same way, stopping a notch short of the edge and never jumping
while the page is still growing.
Typing with a cadence. A field is clicked into for real, then typed key
by key through Input.dispatchKeyEvent with per-character gaps around a
digraph-aware base (common pairs faster, a space and punctuation slower, a
lead-in, occasional longer pauses), a hold on each key, a shift framed around
a capital. The cadence is seeded, so a persona keeps one. A typing step never
presses Enter; press sends a named key (Enter, Tab, arrows, Shift+key)
through the same door.
A step waits for its own fetches. After a click or key that stays on the page, the driver waits for the fetch and XHR requests the page started since the step began (a short grace to appear, a 5 s cap), then for a quiet DOM. A busy page with a long poll no longer costs every step three seconds, and a slow result is waited for rather than read half drawn. A load still waits for the whole network to go quiet.
Submits need approval. A step that would send a form, pay, delete or
confirm returns reason: "submits" with the step to approve: a click, Enter
or Space on a submit, ticking or choosing where that sends the form, a button
or clickable box whose words commit (in English and nine other European
languages). The pending step carries expect: the page (with its query), the
control and the form's values when approval was asked for, and it is pressed
only if they are still that. A replay step the model had to find again is
proposed again, never pressed on the old approval.
Variables. %name% placeholders: the model sees names, the page gets
values, and a secret is hidden wherever the page would show it back, inside
any text, from before it is typed. Inside an each step over rows,
%row.field% is the current row's field.
Tabs and downloads. Only a tab the step itself opened (a popup) becomes
current; others are listed. tabs(), switchTab and closeTab move between
them. Downloads land within limits (downloadsDir, size, extensions), and
uploads go through resolveUpload, so a page gets only files the host allows.
What stands in the way
Notices come first. A bot check (Cloudflare, reCAPTCHA, hCaptcha,
DataDome, HUMAN, Arkose, AWS WAF, Kasada), a refusal (access denied, a 403
or 429), a cookie banner, a dialog, or a sign-in page (a password field with
a sign-in page's words) is written at the top of the outline and returned as
notices. The guarding vendors name what they did in headers only they set
(cf-mitigated, x-datadome, x-amzn-waf-action and the like): a challenge
is a bot check, a denial a block with the vendor's name, before any page
text. The pilot stops at a check with reason: "blocked" instead of spending
model calls on it, after one short wait for the kind that clears by itself.
Nothing here gets past a check: it recognises one so a person can be asked,
instead of retrying into a ban.
Popups. A box drawn over the page is found by where it is drawn, not only
by its markup: a fixed box over the centre or a large share of the view, a
sheet up from the bottom, a scrim with a box on it, a <dialog> or popover in
the top layer (a thin header or footer bar is not one). Right before every
step that acts, one call asks the page whether anything appeared over it, or
took the focus into a field, since the reader last looked: if so, nothing is
done and the step returns reason: "popup" with what it is and its controls.
While a step types, keys reach only its element. Only two kinds are closed
without asking, and each is said in StepOutcome.closed: a cookie banner
(its least-tracking answer) and a banner asking to use the site's app (its
own "Not now", never a control that opens a store, installs, signs in,
accepts or buys). Everything else is the agent's.
Covered elements. When a banner or dialog is on top of the element a step means, the step says so within a second and names what is in the way, with its buttons.
Where it may go. urlPolicy: { allow, block } (host globs such as
*.example.com, or address prefixes) is checked before every navigation and
after every step: a redirect, a link or a new tab that ends somewhere
blocked is undone and the step stops with reason: "blocked-url". API calls,
page reads, site maps and every redirect hop of a call keep to it too, and
private, loopback and link-local addresses are refused unless
allowPrivateNetwork says otherwise.
Data
Data, all of it, checked. extract(instruction, { schema }) reads the
whole page, a chunk at a time when it is long, never only the part a look
shows first. The data is checked against the JSON Schema in full and nested,
each problem with its JSON pointer
(/2/price: should be a number, not a string ("£12")). An answer that does
not fit is asked for once more with the problems listed; then it comes back
with ok: false and the problems. A value the page does not show is null,
never made up.
Lists over pages, and where they stop. With a list schema, pages: {
next?, scroll?, max? } follows the next-page control (found by its intent,
then pressed again with no model) or a page that grows as it scrolls, and
stops cleanly when there is no next page or a page repeats. key: ["sku"]
merges rows that repeat across pages and counts duplicates; limit stops
at that many rows. until: { date } cuts a feed at the first row older than
a cutoff (dates written or relative, "3 hours ago"), and seen keys leave
out rows an earlier run read; a result's keys are what to give back next
time, so a daily run reads only what is new.
Recipes: the same list with no model. After the model reads a list, the
rows it gave are found again in the page itself: the element that repeats
per row, its container, and for each field where it is in a row and how to
read it (by column header, a stable attribute, the label before it, its
role, or its place by tag: never by class name or id). The recipe is kept
only if it reads the model's rows back from the same page. Next time it
reads the list with no model; when it no longer fits, the model reads the
page and a new recipe is learned (relearned).
Routines: record, replay, heal
A recording is a routine (plain JSON, minns.routine/2; /3 when it holds
an api step, /4 when it holds a verify, if, each, grid or doc step; older
files still load). A replay needs no model while the pages match; a moved
element is found by what it is; a renamed one costs one model call from the
step's instruction. Every heal is returned (heals) and the routine comes
back as its next version with its history. keepHeals: "ask" returns it
marked pending, for its owner to keep or not.
A replay stops where the cause is, with no model call. A goto records
where it landed (the path, the title, whether it was a sign-in page); a
replay that opens another path or an error status stops as
landed-elsewhere, and a page that asks for a sign-in the routine did not
record stops as sign-in, before any heal. pilot.verify(text, { absent })
records a check that the page shows, or no longer shows, some words, and a
replay stops as verify when it does not.
Popups in a replay. An act step taken right after a popup stopped or covered a step, on a control that dismisses that popup, is recorded as closing it: on replay it is skipped when the popup is not open, and a popup that meets another step is closed with it and the step taken again.
Flow. if steps take one of two lists of steps on the same check; each
steps take their steps once per item of a list variable, over what an
earlier step read ("@name"), or over the rows of a list on the page:
beginRows()/endRows() record a loop whose list is read by its recipe
each time round, each row's fields given as %row.field%, and a control
recorded inside the first row taken at the same place in each row as a
positional step, never healed.
Proposed variables. withProposedVariables turns the values a recording
typed into variables named for their fields, with the values returned
beside, never kept in the routine.
Memory. memory: { hints, learn } connects any store. When a step is
stuck, hints is asked for notes on the site; they reach the model fenced
as untrusted notes. After the model finds a lost element or a list's recipe
is learned again, learn gets a short lesson about the site: what elements
are called, never a typed value, a secret or the page's text.
Forms as ledgers
pilot.form() reads a form as fields: a stable key
(experience[1].company), the outline id to act on, the label, the kind
(text, textarea, select, combobox, typeahead, radio, checkbox, date, file,
richtext, slider), required, the value, and what the page says is wrong with
it. Custom widgets count; so do fields in iframes (cross-site too) and open
shadow roots. One scan per frame: a 60-field form reads in about 60 ms. A
frame that cannot be read is said in errors, never taken as empty.
pilot.fillForm(answers) fills it as a person would: a typeahead by typing
enough to bring up the option and clicking the one that is the answer
(exact, else starting with it, else containing it; never a near miss), a
select by its words then its value, a date typed in the field's format else
picked on its calendar, an editor by inserting text, a drop zone by dropping
the file, a slider by value or arrow keys. Every write is read back. A value
the page refuses is tried once more with its error in mind. Later rounds fill
fields an answer revealed and press "Add another" for answers that need
another copy of a section. It never submits: no Enter, no submit control,
and the page's own attempts to send the form while it is filled are stopped
and said. pilot.reviewForm() returns every page of a wizard the run has
been through, for the person approving the submit.
Applicant tracking systems. detectAts tells Greenhouse, Ashby, Lever,
Workday, iCIMS and SmartRecruiters apart. For Greenhouse and Ashby the
form's published definition is read once and laid over the ledger: every
option of a dropdown the page has not rendered, required flags, and the
questions no field matched yet.
Applications the outline cannot show
A look says in its first line when the page is a known application (Google Sheets, Excel for the web, Airtable, Google Docs, Notion, Confluence, Google Maps; a host adds its own patterns), and the driver works it as a hand would, with no model:
- Spreadsheets.
grid(),readRange("A1:C20")(the name box, the reference, Enter, copy, the clipboard read in our world: exact) andwriteCell("B2", value)(one cell, typed and committed with Enter, read back). A write of one cell is a write, not a submit; anything larger is a person's to approve. - A sheet understood, with no model.
readSheet()copies the sheet whole (1000 rows by 52 columns by default, trailing blanks cut, the clipboard's HTML read for bold and fill) and finds its tables by shape (detectTables): dense blocks parted by empty rows and columns, a header row told by the type shift under it (text over numbers, dates or yes/no) or by bold, a title just above, each column typed by its majority. The result describes each table in one line, so the agent knows what is there before it reads or fills anything. - extract on a sheet.
extract(instruction, { schema })on a spreadsheet page reads the tables, not the toolbar: the table the instruction names (or whose headers place the most fields), each field placed on a column by its header, and only the fields no header places put to the model over a compressed view of the sheet in the SheetCompressor manner (the header, the first and last rows, the rows where a column's values change kind, the rest counted, every column summarised by type, fill and range or distinct values). The model answers with column letters, never copies cells, so the rows are the sheet's exactly, typed by the schema. The mapping is kept as a grid recipe (gridRecipe,minns.grid/1) anchored to the header row: a replay finds the table again when rows are added above or columns move, and reads it with no model; a recipe that stops fitting is relearned and recorded as a heal. One value off a sheet (an object schema) goes to the model with the tables compressed. - Filling a sheet, approved.
fillCells([{ ref, value }, ...])writes several cells, each typed and read back. One cell is a write. More change the person's document in many places at once, so withoutapprovednothing is written: the result asks for approval and carries what the cells hold now; the approved call, given that back, writes only while they still hold it (changedotherwise). Recorded as a grid fill step that a replay takes only when approved, as a submit; a secret value is typed and never shown. At most 200 cells a fill. - Documents.
doc(),readDocument()(select all, copy, text and HTML off the clipboard, the selection let go),findText(text)andinsertText(text, { after, replace })(at the end, after some words, or in place of one line of at most 200 characters, typed with the hand).
The live selectors of each application are proved by the reliability job; a miss is said, never guessed.
The site's own API
Most lists a page shows came from a JSON call the page made.
Capturing. startCapture() records the network: documents, XHR and fetch
from the page and every frame, who sent each request (the script frames on
the stack, heard through DevTools), the scripts the page loaded as text, and
the frames of WebSockets and event streams. Off until started, and nothing
listens while off. stopCapture() returns it with every secret redacted:
Authorization, Cookie, API keys, CSRF tokens, JWT and key shapes,
secret-named fields, card details, and anything typed as a secret. A
redacted value becomes a marker made from its hash, so the same token is
still recognisable where it came from.
Learning. learnApi({ target: { schema, sampleRows, intent } }) ranks
every response that holds the rows the browser read, and beside them what
the site publishes without a watched request: a feed the page declares (RSS,
Atom, JSON Feed), the JSON view a known platform gives of the very page
(Shopify, WooCommerce, WordPress, Squarespace, Discourse, Reddit, Next.js,
Nuxt), the listing endpoints the site's own code names or its OpenAPI
description lists (GET only, never a path whose words act), and the JSON
answers the application cached for itself in Cache Storage. Every value the
request needs is traced to its source: the page's meta tags and state
(__NEXT_DATA__, WIZ_global_data), local storage, IndexedDB, a cookie, an
earlier response, or a constant in the script that sent the request (a key,
a persisted GraphQL query's hash), kept by the text before it and its shape
so a new build is read again. Form bodies are traced value by value. The
recipe keeps who sent the request, named by the site's source map when it
ships one. A recipe (minns.api/1) is kept only after a direct call gives
back the browser's rows.
Calling. callApi(recipe, inputs, { schema, key, limit, maxPages,
validators }) calls the API from our world of the tab on the site, so the
request carries the browser's own TLS handshake, HTTP/2 connection, cookies,
client hints and referer, the same the page's own call had. Every request
of the tab is paused for the call, so each redirect hop and its preflight is
checked against the URL policy before it is sent, and secret headers are
dropped on a hop to another origin. With no tab on the site, the page the
recipe was learned on is opened first. Tokens are read fresh, every page of
the list is read, a 429 backs the host off with its Retry-After, and the
rows are checked against the schema. Given the last call's validators, a
list the site says is unchanged (304) costs one request and no rows. A
failure says its kind: auth (signs in again with relogin and tries once
more), changed, rate, network, blocked-url.
await pilot.startCapture();
await pilot.goto("https://shop.example/products");
const rows = (await pilot.extract("the products", { schema })).data;
await pilot.stopCapture();
const [recipe] = await pilot.learnApi({ target: { schema, sampleRows: rows, intent: "the products" } });
const all = await pilot.callApi(recipe, {}, { schema, maxPages: 20 });
const routine = withApiStep(recorded, recipe, [0, 1]); // the API first, those browser steps as its fallbackA routine's api step is tried first and its browser steps are skipped when
it answers. When the API has changed, the browser steps read the data
instead while the network is recorded, the API is learned again from what
they read, and the replay returns a heal of kind "api".
Site maps. siteMap(origin, { query }) lists a site's pages without
browsing: its sitemap, its llms.txt, or a shallow crawl, ranked by BM25
against the query. The page's own tools. A page that offers tools to
agents (WebMCP) has them named on the first line of the outline;
callPageTool(name, args) calls one, and one not marked read-only needs
allowSubmit.
The wire below DevTools
network/netlog.ts reads Chromium's own NetLog (a browser started with
--log-net-log), whole or cut short, into a redacted summary: every request
with its status, protocol and timing, the TLS version, cipher and ALPN of
the connection it went over, the net error a request died of before any
HTTP answer, errors by host. The reliability eval runs a job that failed as
blocked or unreachable once more with the log on and keeps the summary in
the record, so a failure reads as "TLS 1.3, h2, a challenge header" or "a
connection reset before any answer". The network fingerprint eval checks
that an API call made through the transport shows the same sorted JA4 and
HTTP/2 settings as a navigation: the request is the browser's, not a
client's standing in for it.
Safety that holds everywhere
- A secret's value never reaches a prompt, an outline, a journal or a
routine. A payment card detail is one whoever typed it: a card field's
value and any card number show as
(hidden)in an outline and a ledger and as a marker in a capture, and nothing here types one on its own. - Page text is data, inside
<page_content>in prompts and quoted in notices and errors. - Notices recognise a bot check so a person is asked; nothing solves, hides from or retries into one. A cookie banner is only ever answered with its least-tracking button.
- Anything that submits needs
allowSubmit; an approved submit also checks the page, the control and the form's values. fillFormnever submits; every write is read back; an answer with no field is said, never put in a near one.- The driver leaves no footprint a page script can see.
- Nothing leaves a capture unredacted, and no recipe holds a secret: a token is an input filled from its source at call time. A recipe is kept only after a direct call gives back the rows the browser read. Discovery probes are GET only, on the site, and never a path whose words act.
API
| | |
|---|---|
| new PageDriver(page, opts?) | look({ part, full, expand, screenshot, forAgent }), perform(step), inspect(ref), tabs(), form(), fillForm(), startCapture(opts), stopCapture(), callApi(recipe, inputs, opts), readUrl(url, opts), readAppCaches(), pageTools(), callPageTool(name, args, opts), grid(), readRange(ref), writeCell(ref, value), readSheet({ maxRows, maxCols }), doc(), readDocument(), findText(text), insertText(text, opts). Options: pointer (the hand model; false keeps Playwright's actions), memory and secrets (where the last look's targets and typed secrets persist), downloadsDir, downloadTimeoutMs, downloadMaxBytes, downloadExtensions, resolveUpload, maxChars (20,000), timeoutMs, maxTabs, apps, answerCookieBanners, closeAppBanners. |
| new BrowserPilot({ driver, model, models?, instructions?, memory?, extractChars?, vision?, urlPolicy?, allowPrivateNetwork?, relogin?, forms? }) | act(instruction, { variables, allowSubmit }), observe(request?), extract(instruction, { schema, pages, key, limit, until, seen, recipe, gridRecipe, as }), goto, verify(text, { absent, waitMs }), pageStep, tabs, switchTab, closeTab, runStep(step, { allowSubmit, expect }), startRecording(), stopRecording(meta), beginRows(), endRows(), replay(routine, variables, { approved, approvals, cellApprovals, keepHeals, runId, lastRows, seen, validators }); form, fillForm, reviewForm, clearForms; startCapture, stopCapture, learnApi, callApi(recipe, inputs, { ..., openSite }), siteMap, pageTools, callPageTool; grid, readRange, writeCell, readSheet, fillCells(writes, { approved, expect }), doc, readDocument, findText, insertText. |
| Network | NetworkRecorder, Redactor, apiCandidates, keepValidated, callApiIn, readUrlIn, browserTransport, clientTransport, sameSite, inferPagination, traceValue, bundleSource, bundleValue, parseFeed, detectPlatform, knownViews, siteViews, endpointLiterals, openApiViews, parseSourceMap, originalPosition, readCacheStorage, readIndexedDb, parseNetLog, summarizeNetLog, parseApiRecipe, withApiStep. |
| Site maps, tools, policy | siteMap, bm25, parseSitemap, parseLlmsTxt; pageToolsInPage; checkUrl, guardDriver. |
| Notices, cards | detectNotices, vendorHeaders, noticeLines; isCardNumber, isCardLabel, hideCardNumbers. |
| captureSnapshot(page, opts?) | The outline and every element's target, without a driver. |
| resolveTarget(snapshot, target, id?) | A recorded element on the page now, or null. |
| Prompts | buildActMessages, buildObserveMessages, buildExtractMessages, buildGridMapMessages and their parsers; wrapPageContent. |
| Routines | parseRoutine (reads /1 to /4), nextVersion, variablesUsed, describeStep, parameterize, withProposedVariables, landedElsewhere, itemsOf. |
| Data | validateSchema, isListSchema, RowCollector, parseWhen, cutRows, recipeInPage, recipeAgrees, parseNumber. |
| Hand | pointerPath, Pointer, keystrokes, keyPress, chord, NAMED_KEYS. |
| Apps | detectApp, GRID_ADAPTERS, DOC_ADAPTERS, isCellRef, parseTsv, sameCellValue, htmlToOutline; detectTables, describeTables, compressGrid, learnGridRecipe, applyGridRecipe, findRecipeTable, parseGridRecipe, parseGridHtml. |
Element methods: click, doubleClick, hover, fill, type, press,
selectOption, check, uncheck, scrollIntoView, upload, drag (onto
another element's id, or by an offset "dx,dy"). Page methods: goto,
back, scroll (with an element: the box it is in), wait (text is looked
for on the whole page; gone waits for it to go), switchTab, closeTab,
clickAt (a point of the view: the control under it gets the ordinary
click, guarded; a canvas or map the press exactly there). Typing never
presses Enter.
Pictures. The outline is the page the model reads; with vision: "auto"
or "always" the pilot also attaches the view's screenshot to its act and
extract questions (Message.images), so a model that reads pictures can
place a point on a map or read a chart, while still answering with the
outline's ids. A text-only model ignores them.
Upgrading: see CHANGELOG.md, which lists what a consumer has to change.
Evaluations
evals/ is not shipped. mind2web/ runs Online-Mind2Web with an agent loop
over the pilot and WebJudge. reliability/ records 19 data jobs once and
replays them daily, with the wire log on a failed job. fanout/ runs one job
over 12 subjects, one at a time and in parallel. network/ is a capture
server of our own that reads the raw ClientHello and the h2 preface, and
diffs a build of the browser (its navigations and its API calls) against a
reference capture of real Chrome. capture/ captures a persona for the
control plane. See evals/README.md.
Requirements
Node 18+, and playwright-core (a peer) for PageDriver and
captureSnapshot. The pilot, the prompts and the routine format need neither.
Development
npm install
npm test # real Chromium: MINNS_CHROMIUM_PATH, Playwright's, or a system one; skipped without
npm run typecheck
npm run buildPushing to main publishes to npm once the suite passes.
The approach to page reading follows Stagehand by Browserbase (MIT); see NOTICE.
