npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@epoch-agent/plugin-web

v0.23.0

Published

Web 抓取插件 — URL 安全校验 + HTML→Markdown + SSRF 防护

Readme

@epoch-agent/plugin-web

网页抓取与搜索插件。两个工具:web_fetch(总是有)和 web_search (有可用后端才注册 —— 缺省状态下就有,见下)。

  • ✅ 做:HTTP GET → HTML→Markdown/Text/HTML、SSRF 校验、逐跳重定向校验、大小限制; web 搜索(native / tavily / brave / searxng / keyless 五个后端)
  • ❌ 不做:不执行 JS、不渲染页面、不带浏览器;不抓搜索引擎的 HTML 结果页 (违反 ToS、随时会碎、要处理验证码 —— 要搜就走后端,不走爬虫)
  • 依赖:只有 node 内置模块。protocol 是 peer, HTML 转换是手写的(无第三方库)

工具参数

{ "url": "https://example.com/doc", "format": "markdown", "timeout": 30 }
{ "query": "node 22 fetch keepalive", "maxResults": 5, "recencyDays": 30 }

两个工具的参数表在 docs/TOOLS.md —— 那份是工具清单 的唯一真源,这里只写这个包自己的实现取舍。

web_search:搜什么,不抓什么

只返回标题 / URL / 摘要,不返回正文。 一次搜索 5 条结果全抓正文轻松几万 token, 而其中多半是模型看一眼摘要就丢掉的。分工是:搜 → 模型挑 → web_fetch 抓那一条。

五个后端,按凭据自动挑

五档(native / tavily / brave / searxng / keyless)各要什么凭据见 docs/TOOLS.md。

不配 search.provider 时按那张表的顺序探测,第一个可用的赢。显式配了哪家就只认哪家 —— 缺 key 时报诊断而不是悄悄换一家:换一家意味着查询发去了另一家公司, 那是用户没同意的事。

一个后端都没有时 web_search 不注册,启动诊断给 skipped 并逐档列出缺什么。 注册一个一调就报错的工具比不注册更糟:模型看见工具表里有它就会反复试,每次烧一轮。

两档不用办账号的(2026-09-15)

上面那句「没有后端就不注册」原来的实际效果是大多数嵌入宿主的用户都没有搜索 —— 三家老后端全要办账号,而宿主多半不会去配。这一档漏配是静默的:模型不知道自己本该 有搜索,就改用浏览器去抓搜索引擎结果页(正是本包第一节说「不做」的那件事)。

  • native(search/native.ts)借 provider.apiKey,让当前 LLM provider 自己的 服务端搜索去干活。首版认 deepseek / anthropic(共用 Anthropic Messages 协议)。 它的结果项没有明文摘要(正文是 encrypted_content),信息走 SearchOutcome.summary。
  • keyless(search/keyless.ts)打 Exa / Parallel 的匿名 MCP 端点, 带游标轮转摊平两家免费额度。查询词会发给第三方,所以它排在最后、且能整档关掉。

两个开关缺省开:search.native: false / search.keyless: false。

SearXNG 默认没开 JSON 输出,实例的 settings.yml 里要有 search: { formats: [html, json] }。没开时它返回 403,我们把这句话写进了错误里 —— 否则用户会以为是自己地址写错了。

搜索结果是风险最高的一种不可信内容

它和抓来的网页一样带 openWorldHint: true(见下一节),但风险更高一档: 网页要等 agent 主动去抓,而搜索结果是攻击者可以针对某个查询词主动投放的 (SEO 投毒 + 提示注入)。所以 core 的 system prompt 在不可信内容策略里单独点了它的名。

结果里的 URL 不自动跟随。模型要抓时走 web_fetch,SSRF 校验和 network 权限判定一道都不会因为「这是搜出来的」而放松。

失败(key 错 / 配额用尽 / 网络不通)返回一条说明原因的失败结果,不抛异常 —— 搜索挂了不该让整轮对话断在这儿。同一个查询在同一会话内缓存,省配额。

SSRF 防护是三层,不是一层

  1. 协议白名单——只放 http / https
  2. 主机名与 IP 字面量黑名单——localhost、云元数据域名、私网段 (10/8、127/8、169.254/16、172.16/12、192.168/16、CGNAT、组播与保留段,IPv6 同理)
  3. DNS 解析后再校验一次——攻击者控制的域名可以解析到 127.0.0.1, 光看主机名字面量挡不住。这是最常见的绕过手法

外加逐跳重定向校验:不用 fetch 的 redirect: 'follow',而是手动跟随、每一跳 重新做完整校验(最多 5 跳)。默认跟随只校验第一个 URL,一个公网地址 302 到 169.254.169.254 就能直接读到云元数据。

其余限制:响应上限 5 MB,Cloudflare 挑战会重试一次。

三个 annotation 不只是元数据

annotations: { readOnlyHint: true, idempotentHint: false, openWorldHint: true }

openWorldHint: true 会让 core 的 tool-executor 把输出包进 <tool_output untrusted="true">。抓来的网页内容作者既不是用户也不是我们,必须标成 「数据」而不是「指令」再回灌上下文。

describeTarget: (args) => args.url 让审批缓存按 URL 记——不给它就只能退回参数 JSON, 「批准抓这个站」下次照样弹窗。

文件

| 文件 | 内容 | | ---------------------------------- | ----------------------------------------------- | | index.ts | webPlugin + createWebSearchTool(条件注册) | | tools/web-fetch.ts | 抓取 + 重定向 + 大小限制 + 重试 | | tools/web-search.ts | 搜索工具:参数钳制 + 会话内缓存 + 失败不炸轮 | | search/provider.ts | 后端接口 + 共用的 HTTP / 错误分类 | | search/registry.ts | 按配置 + 凭据挑后端(挑不到就不注册) | | search/{tavily,brave,searxng}.ts | 三个要办账号的后端各自的请求 / 响应翻译 | | search/native.ts | 借当前 LLM provider 的服务端搜索 | | search/keyless.ts | Exa / Parallel 免费端点环 + 游标轮转 | | search/mcp.ts | MCP tools/call(纯 JSON 和 SSE 两种回法) | | utils/url-safety.ts | 三层 SSRF 校验 | | utils/html-to-markdown.ts | 手写 HTML→Markdown / →Text |

开发

pnpm --filter @epoch-agent/plugin-web test

四个用例文件:ssrf.test.ts、url-safety.test.ts、html-to-markdown.test.ts、 web-search.test.ts。改 SSRF 判定必须先看前两个——里面记的是具体的绕过手法, 不是凑数的。搜索那份全部用假 fetch:真打网络的用例要么依赖别人的 API key、 要么在墙内直接红,两种都不该进 pnpm check。