scuttle-browser
v0.1.1
Published
Browser bridge for LLMs — exposes web pages as annotated accessibility trees via MCP
Maintainers
Readme
Scuttle
Browser bridge for LLMs. Lets AI agents browse the web by converting pages into compact, annotated accessibility trees — no screenshots or raw HTML required.
Scuttle runs as an MCP (Model Context Protocol) server. Any MCP-compatible client (Claude Code, Claude Desktop, Cursor, etc.) can use it to navigate, read, and interact with web pages.
How it works
Instead of feeding raw HTML (too many tokens, too much noise) or screenshots (unreliable for spatial reasoning), Scuttle extracts the browser's accessibility tree — the same semantic structure used by screen readers. It then:
- Prunes invisible and decorative nodes
- Assigns numeric IDs to every interactable element (
[1] button "Submit",[2] textbox "Search") - Labels unnamed elements using a fallback chain: visible text → aria-label → placeholder → nearby heading → positional description
- Accepts actions by element ID —
click(1),type(2, "hello")
The result is a compact, token-efficient representation that LLMs can reason about and act on.
Page: Hacker News
URL: https://news.ycombinator.com/
Interactable elements: 227
────────────────────────────────────────────────────────────
heading "Hacker News"
[1] link "Hacker News"
[2] link "new"
[3] link "past"
...
row
cell "1."
[12] link "Show HN: Something cool"
"Show HN: Something cool"
cell "142 points by user 3 hours ago"
[14] link "user"
[15] link "85 comments"Installation
npm install -g scuttle-browserOr run directly with npx (no install needed):
npx -y scuttle-browserConfiguration
Claude Code
Add to your project's .mcp.json:
{
"mcpServers": {
"scuttle": {
"command": "npx",
"args": ["-y", "scuttle-browser"]
}
}
}Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"scuttle": {
"command": "npx",
"args": ["-y", "scuttle-browser"]
}
}
}Environment variables
| Variable | Default | Description |
|----------|---------|-------------|
| SCUTTLE_HEADLESS | true | Set to false to show the browser window |
| SCUTTLE_VIEWPORT_WIDTH | 1280 | Browser viewport width in pixels |
| SCUTTLE_VIEWPORT_HEIGHT | 720 | Browser viewport height in pixels |
| SCUTTLE_TIMEOUT | 30000 | Navigation timeout in milliseconds |
| SCUTTLE_SETTLE_TIME | 500 | Post-navigation settle time in ms (for SPAs that hydrate after DOM ready) |
Example with visible browser:
{
"mcpServers": {
"scuttle": {
"command": "node",
"args": ["/path/to/scuttle/dist/index.js"],
"env": {
"SCUTTLE_HEADLESS": "false"
}
}
}
}Tools
navigate
Go to a URL.
navigate({ url: "https://example.com" })
→ "Navigated to: Example Domain\nURL: https://example.com/"observe
Get the current page state as an annotated accessibility tree. Every interactable element gets a numeric ID in brackets.
observe()
→ Page: Example Domain
URL: https://example.com/
Interactable elements: 1
────────────────────────────────────────────────────────────
heading "Example Domain"
"This domain is for use in illustrative examples..."
[1] link "More information..."act
Perform actions on the page using element IDs from observe.
| Action | Parameters | Description |
|--------|-----------|-------------|
| click | id | Click an element |
| type | id, text | Clear and type text into an input |
| select | id, text | Select a dropdown option |
| hover | id | Hover over an element |
| scroll | direction ("up" or "down") | Scroll the page |
| key | key (e.g., "Enter", "Tab") | Press a keyboard key |
| wait | ms (max 10000) | Wait for dynamic content |
| back | — | Browser back |
| forward | — | Browser forward |
act({ action: "click", id: 1 })
→ "Clicked link 'More information...'"
act({ action: "type", id: 5, text: "search query" })
→ "Typed 'search query' into textbox 'Search'"
act({ action: "scroll", direction: "down" })
→ "Scrolled down"screenshot
Take a PNG screenshot of the current viewport. Returns a base64-encoded image. Useful when the accessibility tree alone isn't enough to understand the layout.
get_text
Extract all visible text from the page. Best for reading articles, docs, or any text-heavy content without the structural overhead of observe.
Usage pattern
The typical agent loop is:
1. navigate(url) → go to a page
2. observe() → read the page state
3. act(...) → interact with an element
4. observe() → read the updated state
5. repeat 3-4 → until the task is doneArchitecture
┌─────────────┐ ┌──────────────┐ ┌─────────────┐
│ LLM/Agent │◄───►│ Scuttle │◄───►│ Playwright │
│ (MCP client)│ │ (MCP server) │ │ (Chromium) │
└─────────────┘ └──────────────┘ └─────────────┘
│
┌─────┴─────┐
│ │
Accessibility Action
Extractor Executor- Browser Manager — manages Playwright browser lifecycle, navigation, and screenshots
- Accessibility Extractor — snapshots the accessibility tree via CDP, prunes it, assigns IDs, and handles element labeling
- Action Executor — maps element IDs to Playwright locators, executes actions with fallbacks
Element labeling strategy
When elements lack good names (looking at you, <div class="css-1a2b3c">), Scuttle applies a fallback chain:
- Visible text —
[5] button "Add to cart" - Aria attributes —
[6] textbox aria-label="Email address" - Value —
[7] combobox [value: "United States"] - Context —
[8] button (under "Account Settings") - Role fallback —
[9] [unnamed button]
Development
npm run dev # watch mode — recompiles on changes
npm run build # one-time build
npm start # run the MCP serverLicense
MIT
