@pwguler/pi-pengepul-provider
v0.5.3
Published
pi custom provider for pengepul, a relay that pools your Claude/Codex subscriptions. Key and relay base live in auth.json; both native wires (Anthropic Messages, OpenAI Chat Completions) come from one catalog.
Downloads
2,176
Maintainers
Readme
@pwguler/pi-pengepul-provider
A custom provider for pi that connects to pengepul, a local relay that pools your Claude / Codex subscription accounts and serves them over native wire protocols.
pengepul pools several subscription accounts per provider and spreads requests
across them, so pi runs on your subscription instead of a per-token API key.
This extension registers pengepul as a provider so /model shows the models
your relay serves.
Install
From npm:
pi install npm:@pwguler/pi-pengepul-providerOr straight from GitHub (no npm account needed):
pi install git:github.com/pwguler/pi-pengepul-providerPin a release so updates don't move under you:
pi install git:github.com/pwguler/[email protected]To update a git-installed package later:
pi install git:github.com/pwguler/[email protected]Start or reload pi, then select a model with /model. Pengepul models are
prefixed pengepul/<id>. To try it without installing, use
pi -e git:github.com/pwguler/pi-pengepul-provider.
What it does
- Registers the
pengepulprovider against the relay base URL on your credential (http://127.0.0.1:8317by default). - Discovers models from
GET /v1/models, maps each to the right wire:claude-*/anthropic/*andowned_by: anthropic→ Anthropic Messages (POST /v1/messages),gpt-*/o<N>/codex-*and<provider>/<model>→ OpenAI Chat Completions (POST /v1/chat/completions).
- Takes context window, max output, pricing, and image input from what the
relay advertises (pengepul >= 0.6.0 sends
context_window,max_output_tokens,input_modalities,pricing). Fields the relay omits fall back to pi's builtin catalog for the same id, then to family heuristics. - Names the extended thinking levels (
xhigh,max) that a model's own family publishes. An id the catalogs have not caught up with — a point release the relay already serves — takes them from the same family's previous minor, so it does not sit athighuntil the next pi catalog update. Nothing to configure:thinkingLevelMapis resolved per id here. - Publishes Anthropic's cache lifetimes (5 minutes, 1 hour) on the Messages wire, so pi's cache warmer can keep a Claude conversation's entry alive instead of re-billing its whole prefix after five idle minutes. See Prompt cache.
- Caches the catalog in pi's model store, so startup does not wait on the network and a briefly absent relay is covered. The cached models are re-pointed at the relay base configured now, so moving the relay does not leave requests aimed at the old address.
- Reuses pi's built-in stream functions for both wires — no custom transport.
- Registers no commands: pi refreshes the catalog on every startup.
Configuration
Everything lives in ~/.pi/agent/auth.json: the relay's API key and the relay
base it applies to.
{
"pengepul": {
"type": "api_key",
"key": "sk-local-...",
"baseUrl": "http://127.0.0.1:8317"
}
}/login pengepul writes both fields for you — it asks for the key, then for the
relay URL, and defaults the URL to http://127.0.0.1:8317. Editing the file by
hand works the same way; baseUrl may end in /v1 or not.
To reach a relay on another machine, put that machine's address in baseUrl
(and make the relay listen beyond loopback there: host: 0.0.0.0 in
~/.pengepul/config.yaml on the relay, or forward the port over SSH). The key
is the relay's own key — pengepul config api-key prints it.
Optional overrides, for CI or a one-off shell. The first two outrank the
credential: set either and it wins, whatever auth.json holds.
| Setting | Env var | Default |
|---|---|---|
| Relay base URL | PENGEPUL_BASE_URL | http://127.0.0.1:8317, or the credential's baseUrl |
| API key | PENGEPUL_API_KEY | the credential's key |
| Discovery timeout | PENGEPUL_MODELS_TIMEOUT_MS | 10000 |
PENGEPUL_BASE_URL=http://10.0.0.9:8317 PENGEPUL_API_KEY=sk-local-... piDo not set providers.pengepul.baseUrl in models.json. pi applies that value
to every model of the provider, which collapses the two wires onto one URL —
Anthropic Messages traffic would be sent to /v1 and Chat Completions traffic
to /.
Prompt cache
The upstream prompt cache is one entry per account and lives five minutes by
default, so a conversation that waits longer than that pays for its whole prefix
again — on a 700k-token Opus session, dollars per return, logged by the relay as
a cache_write with a near-zero cache_read.
This provider declares Anthropic's lifetimes (5 minutes, 1 hour) on every Messages-wire model, which is what lets pi keep an entry alive at all: pi refreshes an entry by re-sending the prompt with a one-token cap before it expires, and it only warms a model whose lifetime it can resolve. Two pi settings decide when that happens, and neither covers an idle gap out of the box:
| Setting | Where | Effect |
|---|---|---|
| cacheWarming: "idle" | ~/.pi/agent/settings.json | Warms between agent runs. The default, "streaming", stops as soon as a run settles — exactly when a five-minute entry is about to expire. |
| env: { "PI_CACHE_RETENTION": "long" } | the pengepul credential in auth.json | Asks for the 1-hour tier instead of the 5-minute one. The relay forwards the extended TTL to Anthropic. |
{
"pengepul": {
"type": "api_key",
"key": "sk-local-...",
"baseUrl": "http://127.0.0.1:8317",
"env": { "PI_CACHE_RETENTION": "long" }
}
}env is read off the credential by this provider, not by pi core: a provider
that brings its own resolve() has to hand the values back, and pengepul does.
Re-add them after /login pengepul — pi replaces the whole credential on login,
so a key rotation drops the block and, with it, the tier.
Idle warming covers gaps up to 30 minutes, pi's own safety limit, and pi only triggers it when the expected saving clears a threshold of its own — a small prompt is left to expire rather than refreshed. The long tier is what carries a longer break. Both are cheap next to a miss: a refresh is one cache read plus a one-token completion, while the long tier only raises the write premium, on the few hundred tokens a turn actually adds.
/session reports the mode, the refresh cost, and the miss penalty it is
weighing, and showCacheMissNotices in settings.json prints each miss with
its re-billed token count.
Upgrading from 0.4.0 or earlier
This release is breaking: the relay's own ~/.pengepul/config.yaml is no longer
a source. That file belongs to the machine running the relay, and reading it let
a key on that box configure every client sharing it.
- A key that came from
api-keys:in it is no longer picked up. Run/login pengepul, or put it inauth.jsonas"pengepul".key, or exportPENGEPUL_API_KEY.pengepul config api-keyprints it. - The relay base built from its
hostandportis no longer picked up either. SetbaseUrlinauth.json, orPENGEPUL_BASE_URL, if the relay is not onhttp://127.0.0.1:8317. - A pre-0.3
<agent-dir>/pengepul-models.jsonis no longer imported. The catalog now comes from the relay on the next start, so the relay has to be reachable once; pi's model store covers the rest.
Resolution order changed as well, and that is the part to check after an
upgrade: PENGEPUL_BASE_URL and PENGEPUL_API_KEY now win over auth.json,
where they previously lost to it. An exported key from an old shell now
overrides the one in the credential.
Notes
- pengepul >= 0.6.0 advertises per-model context windows, output caps,
modalities, and pricing on
/v1/models, and that is what the provider registers; older relays (or ids the metadata has not reached) fall back to pi's builtin catalog. Your subscription, not a per-token meter, is what pengepul bills against — displayed costs are upstream list prices. - The relay must be running and reachable for discovery to succeed. Without a cached catalog on a first start, pengepul models stay unavailable until a start with the relay up. A relay that answers with 401 leaves the last known catalog registered and logs a warning; discovery does not take pi down with it.
Development
npm install
bun test
npx tsc --noEmit
bun scripts/e2e-live.ts # live e2e against a running relay; sends one tiny completionThe @earendil-works/* copies npm install writes into node_modules are for
tsc and the tests only. At runtime pi serves the extension those modules from
its own bundle (loader.js maps @earendil-works/pi-ai/providers/all to a
virtual module), so the local copies exist to be type-checked against, not to
be the ones in play.
Nothing here can enforce that the two match: the peer ranges are * and the
host's version is not knowable at install time. The lockfile records whichever
version was current when it was last refreshed, so the comparison is a manual
one, worth making whenever a measurement has to be trusted:
pi --version # host pi release
cat node_modules/@earendil-works/pi-ai/package.json # local copyWhen they drift, anything measured from this directory describes a different
model catalog than the running extension sees — the 0.84.4 copy carried 40
:batch entries where 0.85.1 carries 68 — and local verification quietly
disagrees with production.
Releasing
Releases publish from CI on a version tag. Bump, tag, push:
VERSION=0.2.3
npm version "$VERSION" --no-git-tag-version
git add package.json package-lock.json
git commit -m "chore(release): v$VERSION"
git tag "v$VERSION"
git push origin main --tags.github/workflows/publish.yml then runs the full test suite and refuses the
release unless the tag matches package.json, the tagged commit is on main,
and the version is not already on the registry. It publishes with a provenance
attestation linking the tarball to this repository and commit.
Publishing needs a repository secret named NPM_TOKEN, holding a granular
access token:
- Permissions: Read and write. Read-only, and the stage only variant
of read and write, cannot run
npm publish. - Packages and scopes: All packages.
- Bypass two-factor authentication: ticked. A token that prompts for an OTP cannot publish unattended.
- Expiration: whatever you will actually remember to rotate. When it lapses, the release fails at the publish step.
npm removed classic tokens in November 2025, so granular is the only kind that
exists. On the token form, leave Organizations at No access: this package
lives in a user scope (@pwguler), not an organization, so a personal account
has nothing to select there. Choosing Only select packages and scopes instead
of All packages asks for a scope selection that such an account cannot
satisfy.
Trusted publishing (OIDC) is the alternative: no secret to store or rotate, and provenance is generated automatically. It needs a one-time configuration on the package's npm settings page instead of a token.
.github/workflows/ci.yml runs typecheck and tests on pull requests and pushes
to main, and is the same gate the publish job depends on.
License
MIT
