npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@mscststs/pi-provider-gzip

v0.2.0

Published

Pi extension that compresses model request bodies (gzip, Brotli, Zstandard) to cut time-to-first-token on slow LLM gateways.

Readme

English | 简体中文

pi-provider-gzip

一个 pi 扩展,通过 压缩模型请求体,在那些接收大请求体缓慢的 LLM 网关上降低首 token 延迟(TTFT)。默认使用 gzip;Brotli(br)与 Zstandard(zstd) 可按 provider 选用。

before:  2.7 MB request body  ──►  gateway chews on it  ──►  ~55–80 s TTFT
after:   0.75 MB gzip body    ──►  gateway chews on it  ──►  ~12–22 s TTFT
  • 一个开关即可开启或关闭整个功能(PI_GZIP=0)。
  • 默认对所有 provider 生效 —— 无需 allowlist,无需逐 provider 配置。
  • 按 provider 退出 写在 provider 自身的设置里: "compat": { "gzip": false }。
  • 按 provider 选择算法:"compat": { "encoding": "br" }。未配置或值不合法时回退到 gzip。

为什么有用? 许多模型中转站(new-api、one-api、LiteLLM 代理、企业网关……)在派发 请求前会做与接收字节数成正比的工作 —— token 计数、配额预检查、请求体日志、WAF 扫描。这部分开销与模型无关,一旦上下文达到数百 KB 就足以主导 TTFT。压缩 JSON 请求体 可以缩小这些字节。测量数据见 docs/benchmarks.md。

安装

# 从 git
pi install git:github.com/mscststs/pi-provider-gzip

# 从 npm(发布后)
pi install npm:@mscststs/pi-provider-gzip

# 从本地 checkout
pi install /absolute/path/to/pi-provider-gzip

不安装直接试用:

pi -e /absolute/path/to/pi-provider-gzip

安装后,启动一个新会话(或在已有会话中执行 /reload)。

控制项

总开关

| 变量 | 默认值 | 含义 | | --------- | ------ | -------------------------------------- | | PI_GZIP | 1 | 0 表示在当前进程内禁用该扩展。 |

PI_GZIP=0 pi        # 临时禁用全部

高级调优(一般无需修改):

| 变量 | 默认值 | 含义 | | ------------------- | ------ | ------------------------------------------ | | PI_GZIP_MIN_BYTES | 1024 | 只压缩不小于该大小的请求体。 | | PI_GZIP_LEVEL | 6 | zlib 级别,钳制到 0–9。 | | PI_GZIP_DEBUG | 0 | 设为 1 时每次压缩输出日志到 stderr。 |

PI_GZIP_DEBUG=1 pi
# [pi-provider-gzip] relay.example gzip 2708795 -> 753098 bytes (3.6x)
# [pi-provider-gzip] relay.example br   2708795 -> 502144 bytes (5.4x)

级别在所有算法间共享,并映射到各自的取值范围(0–9 → gzip level、Brotli quality 0–11、Zstandard level 1–19)。

按 provider 退出

拒绝 gzip 请求体的 provider 可以在 models.json 中退出:

{
  "providers": {
    "my-relay": {
      "baseUrl": "https://relay.example/v1",
      "api": "openai-completions",
      "apiKey": "...",
      "models": [{ "id": "my-model", "name": "my-model", "contextWindow": 128000 }],
      "compat": { "gzip": false }
    }
  }
}

这也适用于内置 provider —— 不需要 baseUrl,因为单独的 compat 就是合法的 provider 条目:

{
  "providers": {
    "openai": { "compat": { "gzip": false } }
  }
}

compat 是 pi 暴露给扩展的唯一 provider 级字段,因此开关放在这里。 compat: { "gzip": false } 表示退出;其他情况一律启用。

按 provider 选择算法

部分中转站(new-api、one-api 等)也支持解压 Brotli 或 Zstandard 请求体。这两种算法通常比 gzip 再小 15–35%,网关支持时值得开启:

{
  "providers": {
    "my-relay": {
      "baseUrl": "https://relay.example/v1",
      "api": "openai-completions",
      "apiKey": "...",
      "models": [{ "id": "my-model", "name": "my-model", "contextWindow": 128000 }],
      "compat": { "encoding": "br" }
    }
  }
}

| compat.encoding | 发送的 Content-Encoding | | ----------------- | ------------------------- | | 缺省 / gzip | gzip(默认) | | br / brotli | br | | zstd / zstandard | zstd |

其他任何值 —— 未知字符串、空值、非字符串,或当前 Node.js 不提供的算法 —— 一律回退到 gzip。compat.gzip: false 优先于 compat.encoding,会直接为该 provider 关闭压缩。

并非所有中转站都接受 Brotli/Zstandard,而不支持的算法会返回 400 Bad Request (与之相对,未知的值会被忽略)。切换前可用 scripts/probe-encodings.mjs 探测目标 provider 支持哪些算法。

工作原理

在 session_start 时,扩展读取 pi 的实时模型注册表,从所有已配置的 provider 构建一份 host -> encoding 映射,并安装一个进程级的 fetch 拦截器。仅当以下条件全部满足时, 拦截器才会压缩请求:

  1. 方法是 POST;
  2. 目标 host 在映射中(且未被退出);
  3. 请求体是字符串,且长度不小于 PI_GZIP_MIN_BYTES。

随后它会用该 host 的算法压缩请求体(默认 gzip,除非 compat.encoding 另有指定)、设置 Content-Encoding,并删除过时的 Content-Length,让 HTTP 栈重新计算。其他请求一律 原样转发。网关会透明解压;模型看到的 JSON 完全一致。

因为钩子位于传输层,它覆盖 pi 支持的所有基于 fetch 的 API,无论 provider 是内置还是 用户自定义:

| 已覆盖 | 未覆盖 | | --- | --- | | OpenAI Chat Completions / Responses / Azure / Codex (SSE) | Amazon Bedrock(bedrock-converse-stream) | | Anthropic Messages | WebSocket 传输 | | Mistral Conversations | | | Google Generative AI / Vertex | | | pi-messages,以及任意 OpenAI 兼容中转 | |

Bedrock 使用 AWS SDK 的 node:http 传输和 SigV4 签名;WebSocket 传输不经过 fetch。 两者均有意不在范围内。

兼容性

  • 服务端必须接受我们发送的 Content-Encoding 请求体(默认 gzip,或该 provider 的 compat.encoding)。多数网关接受 gzip,接受 Brotli/Zstandard 的较少。若某 provider 以 400 拒绝所选算法,请改用其他 compat.encoding,或用 "compat": { "gzip": false } 关闭压缩。
  • Node: 需要 Node 22.6+(pi 内置兼容的运行时)。Zstandard 需要 Node 22.15+;在更低版本 上 zstd 请求会自动回退到 gzip。

基准测试

在服务大型推理模型、OpenAI 兼容的中转站上,使用真实的 pi 会话测量。完整方法与原始数据 见 docs/benchmarks.md。

| 场景 | 请求体 | 未压缩 TTFT | 压缩后 TTFT | | ---------------- | ------- | ----------- | -------------- | | 全新会话 | 5.7 KB | 1.5–3.4 s | 1.4 s | | 少量历史 | 20.6 KB | 1.6 s | 1.6 s | | 中等历史 | 377 KB | 16.4 s | 3.4 s | | 大量历史 | 1.95 MB | 67.5 s | — | | 超大量历史 | 2.78 MB | 46.8–80.8 s | 11.7–29.4 s |

中转站本身在任意请求体被 gzip 后都返回恒定的 ~0.54 s,而 2 MB 明文请求体则为 48.6 s。

排错

| 现象 | 可能原因 | 解决 | | --- | --- | --- | | 某 provider 返回 400 Bad Request | 它不接受所选 Content-Encoding,或编码名拼写错误 | 设为合法的 "compat": { "encoding": "gzip" },或用 "compat": { "gzip": false } 退出,或 PI_GZIP=0 | | TTFT 没有变化 | 请求体本来就小,或中转站在计费前就解压了 | 用 PI_GZIP_DEBUG=1 检查;收益随请求体大小增长 | | Failed to load extension | Pi 版本不提供此处使用的扩展 API | 升级 pi | | Bedrock 请求从不被压缩 | AWS SDK 传输,不在范围内 | 预期行为 |

开发

npm install        # 开发依赖:pi-coding-agent、@types/node、typescript
npm run check      # 类型检查 + 单元测试
npm test           # 仅 node --test

项目结构

extensions/
  index.ts               # pi 扩展入口(仅做接线)
lib/
  config.ts              # 总开关 + 调优(纯函数)
  encodings.ts           # gzip/br/zstd 编解码 + 回退(纯函数)
  hosts.ts               # 模型注册表 -> host/encoding 映射(纯函数)
  compress-fetch.ts      # fetch 包装器(纯函数)
  interceptor.ts         # 幂等的全局 fetch 安装/卸载
test/                    # node:test 测试套件
docs/benchmarks.md       # 测量记录
scripts/probe-encodings.mjs  # 探测 provider 支持哪些压缩算法

许可证

MIT