@bacnh85/pi-sub
v0.1.52
Published
Pi extension showing subscription usage for supported model providers.
Maintainers
Readme
@bacnh85/pi-sub
Pi extension that shows subscription usage for the currently selected supported model provider.
Supports OpenAI Codex (openai-codex) with live usage windows from ChatGPT's usage endpoint, OpenCode Go (opencode-go) with rolling/weekly/monthly usage windows from Zen's GET /zen/go/v1/usage endpoint plus session cost tracking, and Z.ai GLM Coding Plan — both the international (zai) and China (zai-coding-cn, open.bigmodel.cn) endpoints — with quota monitoring. Also tracks Router (pi-router, router provider) with response-speed tracking and usage windows via yardmaster's general GET /v1/usage API (with OmniRoute's om-usage as fallback), and Command Code (commandcode) 5-hour/weekly windows and monthly credit balance. Displays a subscription footer status after Pi's built-in status/token usage line.
Install
pi install npm:@bacnh85/pi-subWhat it shows
OpenAI Codex
The footer status appears after Pi's built-in status/token usage line and includes:
- active account email/account label;
- subscription plan, such as
Plusin/subdetails; - 5-hour remaining quota and reset countdown;
- weekly remaining quota and reset countdown.
Example subscription line:
([email protected]) R:15%/2H W:20%/3D 42 tok/sOpenCode Go
OpenCode Go reads the rolling (5-hour), weekly, and monthly usage windows from Zen's GET /zen/go/v1/usage endpoint using the stored API key. The footer shows the active account/key label, remaining quota per window, accumulated session cost, and last response speed:
(OpenCode Go key#1a2b3c4d) R:97%/3H W:61%/2D M:20%/5D $0.23 42 tok/sIf the auth entry has an accountId but no API key, the footer falls back to the account label, session cost, and speed (no usage API is called).
Z.ai
Z.ai (GLM Coding Plan) shows the active account/key label, 5-hour rolling and weekly remaining quota with reset countdowns, and last response speed:
(Z.ai key#1a2b3c4d) R:55%/2H W:80%/3D 42 tok/sZ.ai Coding Plan (China)
The built-in zai-coding-cn provider targets the domestic BigModel endpoint (open.bigmodel.cn) and returns the same GLM Coding Plan quota format as the international zai provider, so the footer and /sub detail behave identically, distinguished only by the Z.ai (CN) label:
(Z.ai (CN) key#1a2b3c4d) R:55%/2H W:80%/3D 42 tok/sZ.ai via Anthropic endpoint
The zai-anthropic provider (registered by pi-model-tools, GLM through api.z.ai/api/anthropic) is tracked the same way — same api.z.ai quota monitor as the international zai provider, keyed by the auth.json zai-anthropic credential, labeled Z.ai (Anthropic):
(Z.ai (Anthropic) key#1a2b3c4d) R:55%/2H W:80%/3D 42 tok/sRouter (pi-router — formerly 9router)
For yardmaster instances the footer shows real usage via the general JSON
usage API: GET <baseUrl>/usage?provider=<prefix> with the router API key
(auth.json router credential from /login router, or ROUTER_API_KEY env).
<prefix> is the upstream routing prefix of the selected model
(zai/glm-5.3-flash → zai, cmd/... → command code, ds/... → deepseek);
unknown prefixes fall back to the aggregate report. Unknown/older routers 404
and pi-sub automatically falls back to the OmniRoute text flow below.
Requirements: the API key must hold the usage permission (yardmaster dashboard
→ Endpoints & keys → usage api → "allowed"; keys allow it by default).
Router · zai R:96%/2H W:80%/1D 145 tok/sCredit-based upstreams (DeepSeek) surface their raw balance the same way
(shown as M:$X.XX) — no manage scope needed.
For OmniRoute instances the footer shows real usage: GET <origin>/api/usage/om-usage
with the router API key (auth.json router credential from /login router, or
ROUTER_API_KEY env) returns the per-key report — Personal quota (daily/weekly
USD budgets) and Provider quota (session/weekly connection windows) — rendered
as R:/W: remaining-percent windows:
Router usage R:80% W:28% 145 tok/sRequirements:
- The router instance must be OmniRoute (other routers 404 → endpoint-only fallback).
- The API key must have the usage command enabled in the OmniRoute dashboard (API Keys → the key → enable "usage command"); the footer shows a hint when it's off.
The endpoint URL is read from ~/.pi/agent/settings.json (router.baseUrl), env ROUTER_BASE_URL overrides:
Router (172.30.55.22:20128) 145 tok/sRaw USD balance (credit-based upstreams)
Credit-based upstreams (e.g. DeepSeek) only appear on the om-usage report as
meaningless normalized percentages. When the report has no usable windows,
pi-sub additionally queries OmniRoute's management usage API for the raw USD
balance (shown as M:$X.XX): it discovers the connection id via the
key-authable GET /api/v1/me/status, then reads GET /api/usage/<connectionId>
(quotas.credits_usd.remaining).
The management call needs a manage-scope credential: the router API key itself works when it holds the manage scope (OmniRoute dashboard → API Keys), or set one of these env vars to override it:
ROUTER_MGMT_TOKEN— manage-scope API key oroma_CLI tokenOMNIROUTE_MGMT_TOKEN— legacy alias, checked second
Without either, the router API key from auth.json is used as-is; if it lacks the manage scope the balance fetch simply returns nothing.
Command Code
Command Code exposes live usage windows via its /alpha/billing/credits endpoint (same Provider API key used for /provider/v1 models — no cookies). The footer shows the active account/key label, 5-hour and weekly remaining windows with reset countdowns, and last response speed:
(Command Code key#1a2b3c4d) R:99%/4H W:99%/6D M:$69.99 42 tok/sThe M:$X.XX segment is the monthly credit balance (the plan's remaining
monthly allowance in USD). The /sub detail view also shows a
Monthly: $X remaining line.
Tokens per second
pi-sub tracks each response's tokens-per-second (tok/s) speed by measuring the time from provider request to message completion against the response's output token count. The last response's speed is shown in the footer next to usage data. The /sub detail view shows both the last response speed and the session-wide average.
The tok/s speed line is shown for all providers — including ones without a
subscription adapter, such as ollama and other OpenAI-compatible local
providers. For those, the footer shows only the speed:
145 tok/sand /sub reports the provider/model and speed instead of usage windows.
pi-sub still does not refresh subscription data for unsupported providers
(there is nothing to fetch).
Commands
| Command | Description |
| --- | --- |
| /sub | Show detailed subscription usage for the current supported provider. |
| /sub status | Same as /sub. |
| /sub refresh | Force a usage refresh, then show details. |
| /context | Show a context-window breakdown: total/window/% with a bar, per-section system-prompt costs, tool-schema cost by package, memory files, skills, messages, reserved budget, and free space (TUI only). |
/context prints its panel inline in the transcript (above the input, like OMP's),
so it stays visible while you keep working instead of vanishing on the next keypress.
It renders a 4×10 waffle grid where
⛁ is used context, ⛶ is free space and ⛝ is the autocompact buffer, followed
by a disjoint per-category breakdown. Totals come from Pi's own
getContextUsage(); category splits use the SDK's chars/4 estimate, and the
buffer/free rows are disjoint (free space excludes the buffer, matching Pi's
trigger tokens > contextWindow − reserveTokens):
Context Usage
⛁⛁⛁⛁⛁⛁⛁⛁⛁⛁ GLM-5.3 (197K context)
⛶⛶⛶⛶⛶⛶⛶⛶⛶⛶ glm-5.3[197K]
⛶⛶⛶⛶⛶⛶⛶⛶⛶⛶ 41K/197K tokens (20.9%)
⛶⛶⛶⛶⛶⛶⛶⛝⛝⛝ Estimated usage by category
⛁ System prompt: 12K tokens (6.3%)
⛁ System tools: 4.3K tokens (2.2%)
⛁ System context: 2.1K tokens (1.1%)
⛁ Skills: 3.8K tokens (1.9%)
⛁ Messages: 18K tokens (9.4%)
⛶ Free space: 139K tokens (70.8%)
⛝ Autocompact buffer: 16K tokens (8.3%)
Prompt sections (18K):
rules: 7.5K
tools: 4.0K
skills: 3.8K
project_context: 2.1K
docs: 600
preamble: 300
cwd: 9
Tools (54):
tool0: 80
tool1: 80
tool2: 80
tool3: 80
tool4: 80
pi: 2.7K · 34 tools
pi-web: 1.6K · 20 tools
Recommendations:
- Skills section is 3.8K tokens — disable-model-invocation on reference-only skills.Categories are disjoint — context files and skills are split out of "System
prompt", so the rows add up to the window. Prompt rows are measured character
counts; Messages is the residual (used − prompt rows), because the chars/4
estimate also counts thinking blocks and cache-unwritten content and so
over-reports real sessions by 24–40% (measured). That keeps every row ≤ the
headline total and makes the whole panel reconcile. The buffer reads disabled when
compaction.enabled is false, and honors
compaction.modelOverrides["provider/id"].reserveTokens / compaction.reserveTokens
from settings.json (the same resolution Pi uses), so the number shown is the
number Pi will actually trigger at. Slices worth ≥0.5% always get at least one
grid cell, so a fresh session still shows its weight visually.
When Pi OpenAI Codex auth is available, /sub shows the active account usage and speed:
Provider: Codex · Model: o4-mini · Fetched: 14:23
Session cost: $0.12
Last response: 42 tok/s · Session avg: 39 tok/s
ACCOUNT PLAN ROLLING WEEKLY LAST ACTIVITY
* [email protected] Plus 15%/2H 20%/3D NowFor OpenCode Go, /sub shows the provider/model, active account/key label, rolling/weekly/monthly windows, session cost, and speed:
Provider: OpenCode Go · Model: kimi-k2.6 · Fetched: 14:23
Session cost: $0.23
Last response: 42 tok/s · Session avg: 39 tok/s
ACCOUNT PLAN ROLLING WEEKLY MONTHLY LAST ACTIVITY
------------------------------------------------------------------------------
* OpenCode Go key#1a2b3c4d Go 97%/3H 61%/2D 20%/5D NowFor Command Code, /sub shows the provider/model, active account/key label, rolling windows, and speed:
Provider: Command Code · Model: deepseek/deepseek-v4-flash · Fetched: 14:23
Session cost: $0.05
Last response: 42 tok/s · Session avg: 39 tok/s
ACCOUNT ROLLING WEEKLY LAST ACTIVITY
-----------------------------------------------------------------
* Command Code key#1a2b3c4d 99%/4H 99%/6D Now
Monthly: $69.99 remainingFor Z.ai, /sub shows the rolling and weekly quota windows and speed:
Provider: Z.ai · Model: glm-5.2 · Fetched: 14:23
Session cost: $0.05
Last response: 42 tok/s · Session avg: 39 tok/s
ACCOUNT PLAN ROLLING WEEKLY LAST ACTIVITY
------------------------------------------------------------------------
* Z.ai key#1a2b3c4d Pro 55%/2H 80%/3D NowFor Z.ai Coding Plan (China), the /sub detail is the same, with the account rows showing the Z.ai (CN) provider label (e.g. Z.ai (CN) key#1a2b3c4d in the ACCOUNT column).
Refresh behavior
pi-sub refreshes usage data:
- when a session starts on a supported provider;
- when switching into a supported provider;
- after provider responses, debounced;
- periodically while a supported provider remains active;
- when
/sub refreshis run.
pi-sub reads the openai-codex OAuth entry from Pi's auth file and refreshes live usage directly against ChatGPT's usage endpoint. It does not execute the codex-auth CLI and does not assume a separate Codex CLI installation exists.
Refreshes are cached briefly to avoid excessive usage endpoint calls.
Requirements and troubleshooting
- OpenAI Codex: Pi auth must contain an
openai-codexOAuth entry in~/.pi/agent/auth.jsonor$PI_CODING_AGENT_DIR/auth.json. The entry must includeaccessandaccountIdfields. - OpenCode Go: Pi auth must contain an
opencode-goAPI key entry (via/loginor env var). The entry must have akeyfield or anaccountIdfield. Usage windows come fromhttps://opencode.ai/zen/go/v1/usagewith the stored key; with only anaccountId(no key), the footer falls back to session cost. The Zen credit wallet balance is not exposed by the API. - Command Code: Pi auth must contain a
commandcodeAPI key entry in~/.pi/agent/auth.json(add it withpi /login→ commandcode; there is no env-var fallback). The entry must have akeyfield or anaccountIdfield. Usage is read fromhttps://api.commandcode.ai/alpha/billing/creditswith the same key; 5-hour and weekly windows plus the monthly credit balance are displayed. - Z.ai: Pi auth must contain a
zaientry inauth.jsonwith akeyfield (the same API key used for Z.ai model access via@czottmann/pi-zai-api). The Z.ai provider must be registered (e.g.,pi install npm:@czottmann/pi-zai-api). - Z.ai Coding Plan (China): The built-in
zai-coding-cnprovider targetshttps://open.bigmodel.cn/api/coding/paas/v4. Pi auth must contain azai-coding-cnentry with akeyfield in~/.pi/agent/auth.json(add it withpi /login; there is no env-var fallback). Quota is read from the BigModel endpointhttps://open.bigmodel.cn/api/monitor/usage/quota/limit. - For API-key-only providers, account labels come from stored auth metadata (
email,label,name, oraccountId) when available; otherwisepi-subdisplays a non-secret SHA-256 key fingerprint such asZ.ai key#1a2b3c4d. pi-subredacts auth/token-related errors and never prints credentials.
Changelog
See CHANGELOG.md for release history.
Design notes
The extension is named pi-sub rather than pi-codex-usage so future subscription providers can be added as separate adapters. Supports OpenAI Codex (live usage API), OpenCode Go (live rolling/weekly/monthly windows via /zen/go/v1/usage), the Z.ai GLM Coding Plan (international zai and China zai-coding-cn, which share a quota response format and are served by one parameterized adapter), and Command Code (live 5-hour/weekly windows + monthly balance via /alpha/billing/credits).
