@johnnywu/pi-webfetch
v2.0.0
Published
Fetch web pages and URLs from pi with readable text, Markdown, HTML, or JSON output.
Maintainers
Readme
pi-webfetch
A pi package that adds a webfetch tool for reading web URLs as clean Markdown, HTML, text, or YouTube metadata JSON.
webfetch is optimized for agent use:
- Direct image responses are downloaded as visual input for image-capable models, with inline terminal preview where supported.
- General web pages are cleaned into readable content.
- GitHub URLs are fetched with
ghfor better repository, issue, PR, file, and directory results. - YouTube URLs are fetched with
yt-dlpfor video metadata, transcripts, playlists, and channel listings.
Install
pi install npm:@johnnywu/pi-webfetchOr via local path in ~/.pi/agent/settings.json while developing:
{
"packages": ["~/dev/jwu/pi-webfetch"]
}Requirements
- pi
0.84.x - Node.js
>=22.19.0
Install the optional CLI tools for the URL types you want to support:
| URL type | Required executable | Notes |
|---|---|---|
| Direct images | none | Detected by response Content-Type; JPEG, PNG, WebP, GIF, and AVIF up to 20 MiB |
| General web pages | scrapling | Used for non-GitHub, non-YouTube, non-image URLs |
| GitHub / Gist | gh | Must be installed and authenticated for private or rate-limited content |
| YouTube | yt-dlp | Used for videos, transcripts, playlists, Shorts, and channels |
Defuddle is bundled as an npm dependency and is used by default to improve Markdown output for general web pages.
Tool
webfetch
Fetch and clean an HTTP(S) URL.
| Parameter | Type | Default | Description |
|---|---:|---:|---|
| url | string | required | HTTP(S) URL to inspect and fetch |
| mode | markdown | html | text | json | markdown | Output mode. json is supported for YouTube results. |
Examples:
{ "url": "https://example.com/article" }{ "url": "https://github.com/jwu/pi-webfetch" }{ "url": "https://www.youtube.com/watch?v=PIdETjcXNIk" }{ "url": "https://www.youtube.com/@Brandon-Melville", "mode": "json" }What you get
Direct images
For a general URL whose HTTP response declares image/avif, image/gif, image/jpeg, image/png, or image/webp, webfetch downloads the image directly instead of sending binary data to Scrapling. It returns the image as a Pi image content block for visual analysis and, in terminals that support inline images, displays a preview. Downloads are limited to 20 MiB.
General web pages
Default output is readable Markdown. You can request raw-ish cleaned HTML or plain text with mode: "html" or mode: "text".
GitHub URLs
GitHub URLs are routed through gh, so common GitHub pages return useful CLI/API content instead of noisy browser HTML. Supported URL shapes include:
- repositories
- users
- issues
- pull requests
- releases
- Actions runs
- gists
- files and directories
- commits and other API-backed paths
YouTube URLs
YouTube URLs are routed through yt-dlp.
Supported inputs include:
- video URLs
youtu.beshort links- playlists
- channel handles, for example
https://www.youtube.com/@name - channel Videos / Shorts / Streams tabs
For videos, webfetch returns metadata and tries to include a transcript. Missing transcripts do not fail the request.
For playlists, webfetch returns a flat list of entries.
For channel root URLs, webfetch expands available Videos, Shorts, and Streams tabs and merges their entries into one channel result.
mode: "json" returns a curated stable JSON shape for YouTube video, playlist, or channel data.
Configuration
Add webfetch settings to .pi/settings.json (project) or ~/.pi/agent/settings.json (global):
{
"webfetch": {
"useDefuddle": true,
"qualityJudge": false,
"qualityJudgeModel": "google/gemini-2.5-flash",
"qualityJudgeThinkLevel": "off"
}
}Project settings override global settings after the project has been trusted by pi.
Settings
| Setting | Default | Description |
|---|---:|---|
| webfetch.useDefuddle | true | Use Defuddle to convert cleaned HTML to Markdown for general web pages. Set false to use Scrapling Markdown directly. |
| webfetch.qualityJudge | false | Ask a model to reject unusable fetched Markdown, such as boilerplate, captcha/challenge pages, or unrelated content. |
| webfetch.qualityJudgeModel | current pi model | Optional judge model in provider/model form. |
| webfetch.qualityJudgeThinkLevel | off | Optional judge thinking level: off, minimal, low, medium, high, xhigh, or max. |
Output behavior
- Only
http://andhttps://URLs are accepted. - Missing CLI executables return a friendly failed tool result.
- Output is truncated with pi's standard limits: 2000 lines or 50 KiB, whichever is hit first.
- If output is truncated, the full extracted content is saved to a temp file and the path is included in the result.
Internal docs
Implementation details are documented in:
Development
# Install dependencies
bun install
# Run tests
bun test
# Type check
bun run typecheck
# Format
bun run format
# Release (local, requires GH_TOKEN and NPM_TOKEN)
bun run releaseThis project uses semantic-release with conventional commits.
License
MIT
