@zaphodis42/quota-axi
v0.1.4
Published
AXI CLI that reports local agent-provider quota windows without routing or credential minting
Maintainers
Readme
Quota CLI for agents - designed with AXI (Agent eXperience Interface).
Agents need quota state before they choose where work can safely run. Vendor dashboards are not shaped for shell automation, and local CLIs expose different windows, resets, and auth sources.
quota-axi reports local Claude, Codex, Cursor, GitHub Copilot, Grok, Kimi, Z.AI, Z.ai Coding Plan, Alibaba, OpenCode Go, Antigravity (agy), Command Code, MiniMax, MiMo, DeepSeek, OpenRouter, ElevenLabs, Devin, and Muse quota windows in one AXI-shaped call.
It is data only: it never routes, recommends a provider, model, harness, credential, or route, proxies, intercepts, logs in, imports browser cookies, or mints or rotates a credential. Muse's only quota read also issues an API key server-side, which quota-axi discards unread and rate-bounds (Muse provider notes). When the same stored access token is expired, carries a refresh token, and is definitively rejected, quota-axi may delegate renewal to that vendor's own non-interactive CLI command and re-read the result (Delegated credential refresh). Default output has no ordering preference. The opt-in models --sort runway surface applies only its documented deterministic comparator to quota evidence, preserves all evidence and explicit ties, and is not a recommendation. It publishes one derived per-scope comparative selection signal, selection, as data computed from figures it already reports; the consumer, not quota-axi, does any routing or ranking with it.
- Official sources - quota-axi reads local provider auth sources and calls first-party quota, usage, billing, entitlement, local loopback, or read-only credential-liveness endpoints used by the local agents, with read-only CLI probes where applicable. Vendor-command boundaries and the explicit inference exception are documented under Safety guarantees.
- Local first - quota and auth reports run on the machine that holds the credentials; their network calls go to first-party provider endpoints, never a third-party relay.
The separate
updatecommand contacts npm only when the user runs it. - Token efficient - default stdout is compact TOON so agents spend fewer tokens parsing quota state, with
--jsonavailable when a caller needs the normalized model.
Quick Start
Credential-source note: Claude Code, the Cursor CLI (cursor-agent), Copilot CLI, and Muse Code keep live tokens in native secure stores on supported platforms. Linux cursor-agent stores its access token in ~/.config/cursor/auth.json (or the XDG/$CURSOR_CLI_CONFIG override); Copilot CLI uses Windows Credential Manager on Windows and the macOS Keychain on macOS; Muse Code on macOS stores the OAuth token in the login Keychain item its auth.json storage: "keychain" records point to.
quota-axi does not read native secure-store values until the user grants permission, so Claude quota, CLI-only Cursor auth, Copilot CLI auth, or Muse quota can stay stale or appear unavailable when no other usable credential exists. On Linux it reads only the Cursor auth file's accessToken and never its refresh token.
Run quota-axi --allow-keychain-prompt once and approve macOS Keychain access with "Always Allow"; on Windows the same flag permits the Copilot CLI Credential Manager read without an OS prompt.
After a successful read, future non-interactive quota calls reuse the corresponding account-scoped grant without requiring the flag. Claude grants are also profile-scoped; legacy Claude markers created before account-pinned lookup are not reused.
$ npx -y @zaphodis42/quota-axi
bin: ~/.npm/_npx/.../quota-axi
description: Report local agent-provider quota windows for routing-aware agents
generatedAt: "2026-03-15T16:42:00.000Z"
quota[12]{provider,scope,effectivePercentRemaining,spendPriority,runway,confidence,limitedBy,resetsAt}:
claude,all_models,64,-0.3798,projected_exhaustion,established,seven_day,"2026-03-20T17:59:45.600Z"
claude,seven_day_opus,64,0.3218,projected_exhaustion,established,seven_day,"2026-03-20T17:59:45.600Z"
claude,"model:fable",64,-0.0932,projected_exhaustion,established,seven_day,"2026-03-20T17:59:45.600Z"
codex,all_models,47,-0.2383,projected_exhaustion,established,weekly,"2026-03-19T09:54:28.800Z"
codex,"model:gpt-5.1-codex",47,-0.1973,projected_exhaustion,established,weekly,"2026-03-19T09:54:28.800Z"
cursor,all_models,72,1.4067,through_reset,established,included_usage,"2026-04-01T00:00:00.000Z"
grok,all_products,67,0.5778,through_reset,established,week,"2026-04-01T00:00:00.000Z"
kimi,all_models,74,0.2484,through_reset,established,weekly,"2026-03-20T12:17:02.400Z"
zai,all_models,50,-1.0046,projected_exhaustion,established,weekly,"2026-03-20T16:42:00.000Z"
zai,tools,100,unknown,unknown,unknown,mcp_month,"2026-04-01T00:00:00.000Z"
agy,gemini,88,unknown,unknown,unknown,gemini_weekly,"2026-03-20T00:00:00.000Z"
agy,claude_gpt,90,unknown,unknown,unknown,claude_gpt_weekly,"2026-03-21T00:00:00.000Z"
exhaustion[6]{provider,scope,usableRunwaySeconds,projectedExhaustedAt,limitingWindowId}:
claude,all_models,298906,"2026-03-19T03:43:45.600Z",seven_day
claude,seven_day_opus,298906,"2026-03-19T03:43:45.600Z",seven_day
claude,"model:fable",298906,"2026-03-19T03:43:45.600Z",seven_day
codex,all_models,10365,"2026-03-15T19:34:45.428Z",five_hour
codex,"model:gpt-5.1-codex",10365,"2026-03-15T19:34:45.428Z",five_hour
zai,all_models,172800,"2026-03-17T16:42:00.000Z",weekly
attention[4]{provider,scope,kind,detail,remedy}:
copilot,all,unresolved_windows,chat + premium_interactions,none
zai,tools,unmeasurable,"mcp_month blocks runway + spendPriority",none
agy,gemini,unmeasurable,"gemini_5h + gemini_weekly blocks runway + spendPriority",none
agy,claude_gpt,unmeasurable,"claude_gpt_5h + claude_gpt_weekly blocks runway + spendPriority",none
help[1]:
Run `quota-axi --full` for windows, pace, reserve, and account evidenceDefault TOON is decision-shaped: quota[] carries one fully populated row per measurable scope, and the sparse exhaustion[] and attention[] blocks carry the finite-runway and non-nominal facts. See Default report blocks.
--json emits the normalized model instead. Derivation inputs are demoted to --full; see Output tiers.
$ quota-axi --provider claude --json
{
"generatedAt": "2026-03-15T16:42:00.000Z",
"schemaVersion": 5,
"providers": [
{
"provider": "claude",
"accountKeys": ["default"],
"plan": "pro",
"windows": [
{
"id": "five_hour",
"label": "session",
"kind": "session",
"percentRemaining": 82,
"resetsAt": "2026-03-15T20:10:48.000Z",
"pace": {
"status": "behind",
"reservePercentPoints": 12.4,
"burnMultiple": 0.5921
}
},
{
"id": "seven_day",
"label": "week",
"kind": "weekly",
"percentRemaining": 64,
"resetsAt": "2026-03-20T17:59:45.600Z",
"pace": {
"status": "ahead",
"reservePercentPoints": -8.2,
"burnMultiple": 1.295
}
},
{
"id": "model:fable",
"label": "Fable week",
"kind": "model",
"percentRemaining": 71,
"resetsAt": "2026-03-20T08:25:12.000Z",
"pace": {
"status": "behind",
"reservePercentPoints": 4.5,
"burnMultiple": 0.8657
}
}
],
"state": {
"status": "fresh",
"stale": false
},
"quotaSemantics": {
"status": "known",
"effectiveAvailability": [
{
"scope": "all_models",
"status": "known",
"effectivePercentRemaining": 64,
"boundedBy": [
"five_hour",
"seven_day"
],
"limitingWindowIds": [
"seven_day"
],
"pace": {
"status": "mixed",
"aheadWindowIds": [
"seven_day"
],
"worstReservePercentPoints": -8.2,
"worstReserveWindowId": "seven_day"
},
"runway": {
"status": "projected_exhaustion",
"usableRunwaySeconds": 298906,
"projectedExhaustedAt": "2026-03-19T03:43:45.600Z",
"limitingWindowId": "seven_day",
"projectionConfidence": "established"
},
"selection": {
"status": "known",
"spendPriority": -0.3798
}
},
{
"scope": "model:fable",
"status": "known",
"effectivePercentRemaining": 64,
"boundedBy": [
"five_hour",
"seven_day",
"model:fable"
],
"limitingWindowIds": [
"seven_day"
],
"pace": {
"status": "mixed",
"aheadWindowIds": [
"seven_day"
],
"worstReservePercentPoints": -8.2,
"worstReserveWindowId": "seven_day"
},
"runway": {
"status": "projected_exhaustion",
"usableRunwaySeconds": 298906,
"projectedExhaustedAt": "2026-03-19T03:43:45.600Z",
"limitingWindowId": "seven_day",
"projectionConfidence": "established"
},
"selection": {
"status": "known",
"spendPriority": -0.0932
}
}
]
}
}
]
}$ quota-axi auth
bin: ~/.npm/_npx/.../quota-axi
description: Inspect local quota auth sources without printing secret values
auth[31]{provider,source,path,status,error}:
claude,oauth-file,~/.claude/.credentials.json,available,none
claude,keychain,none,skipped,keychain_prompt_required
codex,auth-json,~/.codex/auth.json,available,none
codex,pi:openai-codex,~/.pi/agent/auth.json,available,none
codex,cli-rpc,~/.local/bin/codex,available,none
cursor,state-vscdb,~/Library/Application Support/Cursor/User/globalStorage/state.vscdb,available,none
cursor,cli-keychain,~/.cursor/cli-config.json,skipped,keychain_prompt_required
copilot,apps-json,~/.config/github-copilot/apps.json,available,none
copilot,copilot-cli:keychain,~/.copilot/config.json,skipped,keychain_prompt_required
copilot,gh:hosts.yml,~/.config/gh/hosts.yml,available,none
grok,auth-json,~/.grok/auth.json,available,none
grok,pi:xai,none,missing,none
kimi,pi:kimi-coding,none,available,none
kimi,kimi-code-cli,none,available,none
zai,pi:zai,~/.pi/agent/auth.json,missing,none
zai,opencode:auth.json,~/.local/share/opencode/auth.json,available,none
agy,loopback,none,available,none
alibaba,bl-cli,none,available,none
opencode-go,opencode:auth.json,~/.local/share/opencode/auth.json,available,none
commandcode,pi:commandcode,~/.pi/agent/auth.json,missing,none
commandcode,env:COMMAND_CODE_API_KEY,none,missing,none
commandcode,env:COMMANDCODE_API_KEY,none,missing,none
commandcode,commandcode-cli,~/.commandcode/auth.json,missing,none
commandcode,omp:commandcode,~/.omp/agent/auth.json,missing,none
minimax,env:MINIMAX_API_KEY,none,missing,none
minimax,pi:minimax,~/.pi/agent/auth.json,available,none
minimax,minimax:config.json,~/.mmx/config.json,missing,none
mimo,env:MIMO_API_KEY,none,available,none
deepseek,env:DEEPSEEK_API_KEY,none,missing,none
deepseek,pi:deepseek,~/.pi/agent/auth.json,available,none
openrouter,env:OPENROUTER_API_KEY,none,missing,none
openrouter,pi:openrouter,~/.pi/agent/auth.json,available,none
help[1]:
Run `quota-axi --allow-keychain-prompt auth` to permit native secure-store accessInstall
quota-axi requires Node.js 22.19 or newer.
Agent skill (recommended)
Install the skill in the Agent Skills format with npx skills:
npx skills add zaphodis42/quota-axi --skill quota-axi -gThe minimal skill points your agent to quota-axi's live CLI guidance through npx -y @zaphodis42/quota-axi, so nothing needs to be installed ahead of time and installed skill copies do not duplicate changing CLI instructions.
The skills CLI's -g option is intended to install for all projects (e.g. ~/.claude/skills/).
It may report PromptScript does not support global skill installation even when the skill was installed for other agents.
Check the relevant agent's skills directory for quota-axi/SKILL.md to confirm.
Drop -g for a per-project install (.claude/skills/), or use npm install -g @zaphodis42/quota-axi or npx -y @zaphodis42/quota-axi for the CLI.
Direct use
npx -y @zaphodis42/quota-axinpm
npm install -g @zaphodis42/quota-axiFrom source
git clone https://github.com/zaphodis42/quota-axi.git
cd quota-axi
pnpm install
pnpm run build
pnpm run devAgent Skill
The npm package includes skills/quota-axi/SKILL.md, the same installable skill recommended above.
It is generated from src/skill.ts as a minimal stub: discovery frontmatter, what quota-axi is, when to reach for it, and pointers to the live CLI (npx -y quota-axi, --help, --json / --full). CLI output is the source of truth. Do not restate CLI instructions or claim that quota-axi recommends or ranks. Update the skill with pnpm run build:skill and verify it with pnpm run build:skill -- --check.
How It Works
┌────────────┐
│ quota-axi │
└─────┬──────┘
▼
┌───────────────┐
│ provider │
│ adapters │
└─────┬─────────┘
▼
┌───────────────┐ ┌──────────────┐
│ local auth or │ ───▶ │ first-party │
│ runtime │ │ APIs/loopback│
└─────┬─────────┘ └──────┬───────┘
▼ ▼
┌───────────────┐ ┌──────────────┐
│ CLI │ ───▶ │ normalized │
│ fallbacks │ │ quota model │
└─────┬─────────┘ └──────┬───────┘
▼ ▼
┌───────────────┐ ┌──────────────┐
│ stale cache │ ◀─── │ TOON/JSON/TUI│
└───────────────┘ └──────────────┘- Live first - provider HTTP calls and Antigravity's structured print command use 15 second timeouts, Codex JSON-RPC and Antigravity loopback reads use shorter per-call timeouts, and stale cache fallback is per provider.
- Host network policy - Claude, Codex, Copilot, Cursor, Grok, Z.AI, Command Code, MiniMax, DeepSeek, OpenRouter, ElevenLabs, Devin, and Muse outbound HTTP calls honor standard
HTTP_PROXY,HTTPS_PROXY, andNO_PROXYenvironment variables (including lowercase forms). This only follows the user's configured egress path; quota-axi does not expose a proxy service or print proxy URLs. New remote fetches go throughsrc/lib/http.tsrather than globalfetch. The proxy path pairs the installedundicibuild'sProxyAgentwith that same build'sfetch, because Node's global fetch only accepts a dispatcher from the undici build Node bundles. - Opt-in secure-store access - macOS Claude, Cursor CLI, Copilot CLI, and Muse Keychain value reads, plus Windows Copilot Credential Manager reads, are skipped on plain calls until
--allow-keychain-promptsucceeds once for that source, then future quota calls reuse the corresponding grant. - Delegated refresh, never minted - when the same stored access token is expired, carries a refresh token, and is definitively rejected, quota-axi may run that vendor CLI's own smallest non-interactive refresh command and re-read the store the CLI rewrote. quota-axi never performs a refresh-token exchange itself. See Delegated credential refresh.
- Partial success is success - one provider can fail while another returns fresh or stale data, and the process still exits 0. Exit code 1 means every provider failed and the report is still rendered. Exit code 2 is a usage error (
VALIDATION_ERROR).quotais the implicit default command: a bare call or a flag-first call is read asquota, whileauth,update, a single-token--help, and a version flag stay with the SDK. Slow-path routing, help, exit framing, andupdatecome fromaxi-sdk-jsrunAxiCli; machine TOON/JSON stays insrc/render.tsand command bodies insrc/commands.ts. A bare-v/-V/--versionis answered byaxi-sdk-js/fast-pathplus the leafsrc/version.ts(node builtins only), which imports the CLI only on the slow path so the provider graph never loads. Keepsrc/version.tsleaf-clean. - No token equivalence - quota-axi does not claim that one provider percentage equals another provider percentage.
CLI Reference
| Command | Description |
| ---------------- | ---------------------------------------------------- |
| quota-axi | Report supported local quota windows |
| auth | Report local auth-source availability, no values |
| models | Join curated model buckets with local quota evidence |
| update | Upgrade quota-axi to the latest published version |
| update --check | Report current vs. latest without installing |
Flags
| Flag | Description |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- |
| --provider claude,codex,cursor,copilot,grok,kimi,zai,agy,zai-coding-plan,alibaba,opencode-go,commandcode,minimax,mimo,deepseek,openrouter,elevenlabs,devin,muse | Scope providers; repeat to union in first-seen order (--provider zai --provider codex equals --provider zai,codex) |
| --json | Emit normalized JSON instead of TOON for quota, auth, or models |
| --full | Include audit and derivation details |
| --tui | Render the live human terminal report instead of TOON (quota only) |
| --refresh 30s\|5m\|1h | Live --tui refresh interval, default 5m (30s-24h) |
| --once | Render one --tui frame and exit instead of staying live |
| --all | Draw not-set-up providers as full --tui cards |
| --allow-keychain-prompt | Permit native secure-store access that could prompt (macOS or Windows) |
| --allow-claude-inference | Spend one bounded native Claude inference to read env-token quota headers |
| --no-credential-refresh | Never run a vendor CLI's own non-interactive credential refresh |
| --max-age 0\|90s\|2m | Opt in to reusing a fresh reading up to this old instead of asking the vendor; 0 always asks |
| --profile-only | Read one explicitly selected Claude or Codex credential file (quota only) |
| --intelligence high\|medium\|low | Filter models by editorial intelligence bucket |
| --sort runway | Explicitly sort models by documented usable-runway evidence |
| -h, --help | Print terse AXI help |
| -v, -V, --version | Print version |
Profile-only quota reads
Profile-only mode is accepted only by quota, requires exactly one --provider selector, and supports only Claude and Codex. Claude requires an explicit nonblank CLAUDE_CONFIG_DIR; Codex requires an explicit nonblank CODEX_HOME. There is no default-location fallback in this mode, and --allow-keychain-prompt and --allow-claude-inference are rejected rather than ignored.
It reads only $CLAUDE_CONFIG_DIR/.credentials.json or $CODEX_HOME/auth.json. It never reads the macOS Keychain or Pi auth, invokes a CLI RPC or other credential fallback, delegates a refresh, or reads, writes, clears, or persists quota cache data. --full --json retains non-secret account identity, the top-level source, and source attempts for provenance. Tokens and credential-file contents remain excluded. Ordinary output remains redacted.
CLAUDE_CONFIG_DIR=/path/to/claude-profile quota-axi --provider claude --profile-only --full --json
CODEX_HOME=/path/to/codex-profile quota-axi --provider codex --profile-only --full --jsonFresh reuse
A burst of back-to-back reads (a dispatcher that takes a fresh reading per decision, a test run) can trip a vendor's usage-endpoint rate limit and leave the provider unmeasured until the limit clears.
Fresh reuse lets such a caller serve the last successful reading instead of asking the vendor again.
It is opt-in: with neither --max-age nor QUOTA_AXI_MAX_AGE, every quota and models read asks the vendor.
--max-age <duration> enables it for one call, and QUOTA_AXI_MAX_AGE=<duration> enables it for a host; the flag wins, so --max-age 0 always asks the vendor.
Both accept 0, bare seconds, or a whole-unit duration up to one hour, and a QUOTA_AXI_MAX_AGE that does not parse fails the read with a validation error instead of being ignored.
Ninety seconds absorbs a burst and moves a five-hour window's elapsed time by half a percent.
A reading is reused only when all of these hold:
- It came from the same credential selection: every environment variable a provider consults to choose a profile, store, credential, CLI, or deployment (
CREDENTIAL_SELECTION_ENVinsrc/lib/reuse-context.ts) has the value it had, plus the same home directory and user. The values are hashed together and never stored. - No local file the reading was derived from has changed. Every file quota-axi opened for that provider, present or absent, is recorded with its identity, size, and modification times (never its contents), so a login that rewrites a credential store, or a Kimi Code configuration that now names another deployment, is read again. A store read through a vendor tool (the macOS Keychain, Cursor's SQLite database, a vendor CLI) is identified by the selection that names it, the same boundary the stale-cache context uses, so an in-place Keychain login switch can be reused for at most
--max-age. Muse is excluded from this path: its macOS credential lives in a Keychain item that is not a traced file, so a login switch cannot be detected here, and--max-agenever serves a Muse reading. The five-minute key-endpoint ledger remains the bound on those vendor calls. - Every lane the provider reported in that reading was a fresh reading with windows, so part of an account-expanded report is never served, and a failed, stale, or window-less lane makes the next read ask the vendor.
- It is younger than
--max-age, and no window has reached its own reportedresetsAt, because a number that has stopped being true is never served (#257).
The cache keeps one slot per provider lane, so two profiles that alternate each read the vendor; each still reuses its own reading between switches.
--full is the audit tier and account identity and source attempts are never cached, so QUOTA_AXI_MAX_AGE does not reach it; it reuses only when --max-age is passed explicitly.
--profile-only never reuses, and readings the cache excludes (Claude native inference and Copilot native secure-store readings) are never reused.
With reuse enabled, a live --tui reuses on its first frame; a scheduled refresh reuses only a reading younger than the refresh interval, so it never repeats its own previous frame and still reads the vendor at any --refresh, unless --max-age is passed explicitly.
Pressing r always reads the vendor, because it is an operator asking for a new reading now.
With reuse enabled, processes that miss the cache together make one vendor read per provider and credential selection. The first creates a lock file under the cache directory, reads the vendor, and caches that provider's reading before releasing it; the others poll the cache and are answered by that reading through the same checks as any reuse. A waiter gives up after 30 seconds and reads the vendor itself, and a lock whose holder has exited, or that is older than two minutes, is taken over, so a crashed or wedged holder never blocks a read. If the holder's reading is not reusable (it failed, is stale, has no windows, or is a reading the cache excludes), nothing is cached and the waiters read the vendor together rather than one at a time, so a burst takes about one vendor round trip beyond the holder's, not one per waiter. Leaders of different providers merge their readings into the one cache file under a short lock of their own, so one leader's write never drops another's reading and leaves its waiters to read the vendor; if that lock is not free within five seconds the write goes ahead unlocked. The lock only saves vendor calls: a lost race costs an extra read, never a wrong one.
A new environment variable a provider reads to choose a credential must be listed in CREDENTIAL_SELECTION_ENV (src/lib/reuse-context.ts) or declared non-selecting in test/reuse-context.test.ts. A new provider file read goes through the src/lib/fs.ts readers or traceInput (src/lib/input-trace.ts). quota-axi's own cache is read untraced, because every write changes it. QUOTA_AXI_SNAPSHOT is the fixture surface on the same path.
A reused reading is a fresh reading, not a stale one: state.status stays fresh, stale stays false, and quota rows, pace, runway, and selection are derived as usual at the report's generatedAt.
It is marked honestly on every surface: --json adds state.reused: true and keeps its state.refreshedAt (the time the vendor answered) in the default tier, default TOON adds a reused attention row naming that time, and the --tui card title says reused 42s.
Reusing never rewrites the cache record, so its age keeps counting from the vendor's answer.
Serving it as fresh is honest because the caller chose the bound: a reused reading is the vendor's own answer at refreshedAt, no older than the --max-age the caller asked for, and never past a window's reset, so a consumer that needs a newer number reads refreshedAt or passes a smaller --max-age.
Snapshot fixtures
QUOTA_AXI_SNAPSHOT=<file> makes quota and models answer every provider from that file instead of any credential or vendor, for tests and fixtures of tools that consume quota-axi output.
The file uses the cache file format ({"schemaVersion": 3, "providers": [...]}), so a cache file captured from a real run is a valid fixture.
Each provider's snapshots are served as reused readings with sourcesTried: ["snapshot"], whatever their age; a provider the file does not name reports unavailable with not_in_snapshot, and one with a window past its own reset reports unavailable with snapshot_expired.
A file that is missing, does not parse as that format, or holds a provider record that does not parse fails with VALIDATION_ERROR (exit 2) rather than reporting that provider as not_in_snapshot.
Nothing is written to the cache, and the variable cannot be combined with --profile-only.
QUOTA_AXI_SNAPSHOT=test/fixtures/quota.json quota-axi --jsonHuman terminal report (--tui)
quota-axi --tui renders the same redacted report as a live human terminal view instead of TOON: a two-up provider card grid with thin quota bars and a ┃ linear-pace marker whenever pace is known. It is presentation only and is not part of the machine-readable contract.
- On an interactive terminal the report stays up and refreshes every 5 minutes until you press
q(or Ctrl+C). Pressrto refresh immediately; the footer lists both controls.--refreshsets the interval (30s-24h) and--oncerenders a single frame. A non-TTY stdout or stdin (pipes, CI, screenshots) always renders one frame and exits. - Every refresh re-runs the same quota read as a bare
quota-axi, including delegated credential refresh when a stored session has expired in the meantime. Runquota-axi --tui --no-credential-refreshto keep the live report strictly read-only. - Live frames paint on the alternate screen and repaint immediately on terminal resize; quitting restores the screen and prints the final frame so the last report stays in scrollback.
- Height comes from the terminal too. When the report is taller than the terminal, the live view windows it instead of letting the alternate screen (which has no scrollback) push the header and first cards out of reach. The viewport accounts for physical rows after terminal-width wrapping: a full-width visible line can consume its own row without wrapping the optional scroll affordance, which is omitted when it cannot fit. At five or more rows the header stays pinned; when there is room, the last row carries a scroll affordance naming how many report lines are above and below. Below five rows the header scrolls with the other report content, and at one row the scroll affordance is omitted so content still remains visible. Use
j/k, the arrow keys,PgUp/PgDn,Space/b, org/Gto move the window.renderQuotaTuilays cards out for the width and returns every line that takes;scrollFrameinsrc/tui-viewport.tswindows that frame to the terminal, pins the first line, and clips oversized lines. The closing hint is composed outsiderenderQuotaTui, and the live paint has no trailing newline, so a frame that exactly fills the terminal does not scroll itself. Fold a provider only on positive evidence of absence. Declare an uncertain skip on the adapter that owns the source (isUncertainSkip) and a sibling-tool login as incidental (incidentalSources); do not special-case those in the TUI. Thetui.showpreference stays on the human path insrc/commands.tsandsrc/tui.tsand never changes the model or machine output. Scrolling clamps at both ends, survives a refresh, and re-clamps on resize; growing the terminal back past the report's height restores the whole frame and the ordinaryPress r to refresh · q to quitfooter.--once, non-TTY output, the final frame echoed on quit, and the TOON and JSON surfaces are all unaffected by terminal height. - Each live card with a combinable bound leads with the effective-availability rollup (min across bounding windows), colored by headroom: >=50% healthy, 20-50% tight, <20% critical. Per-window rows, including per-model breakouts, are the supporting detail.
- Percentages and bars show what is left by default. To have every provider's headline, window rows, and bars show what has been used instead, set
tui.showtousedin the user config file~/.config/quota-axi/config.json($XDG_CONFIG_HOME/quota-axi/config.jsonwhenXDG_CONFIG_HOMEis set):{ "tui": { "show": "used" } }. That file is the only place the preference is read - there is no environment variable or flag for it - and it is read once when--tuistarts, so restart a live report to pick up a change. Only the exact valuesusedandremainingare recognized; a missing, unreadable, or malformed file, or any other value, keeps the default remaining view. The used view rounds the raw consumed percentage independently of the remaining view, so rounded figures need not sum to 100% (48.5% used / 51.5% remaining displays 49% used / 52% remaining). It labels each headline49% used · week. Its bar fill is consumption, the┃marker moves to the elapsed share of the window, and fill running past the marker means burning ahead of the reset clock. Health colors, runway verdicts, used-share rows, raw credit balances, and?rows read the same in both views. It is a human display preference only: quota-axi never reads the file outside--tui, so the default TOON,--json,--full, and the cache are byte-identical whatever it says. - A window with
shareOfis a used-share, not an independent remaining pool: the row printsN% of <parent>(orshare of <parent>whenpercentUsedis absent) instead of a remaining bar or?, so it cannot be read as missing data or as its own headroom. - The headline is labeled with the window it actually is: the minimum across bounding windows always equals at least one named window, so the label names the
limitingWindowIdswindow (week,session,credits) and changes per provider and over time. Tied limiting windows readcredits + grok build, compacting tocredits +2when the names do not fit; a model- or product-scoped headline appends its scope, and any unresolved limiter falls back to the scope wording (all models). - In the default view, the bar fill is current headroom; the
┃marker sits at the binding window'space.timeRemainingPercent, the fill position of exactly linear burn. The headline marker therefore matches the correspondinglimitingWindowIdssub-bar even when another window supplies the finite-runwayempty inverdict. Fill ending left of the marker means burning faster than the reset clock. The marker is omitted when that window's pace is unknown. - Pace is shown by the bar and marker alone, never as a numeric burn multiple. The runway verdict on the headline reads
on pace ✓forthrough_resetandempty in 7h 21mforprojected_exhaustion. Two-up rows keep both card bottoms aligned by padding the shorter card inside its border. The TUI does not display the per-scope selection signal; that signal remains on the JSON and TOON machine surfaces. Those surfaces also keep thethrough_resetvocabulary, while--full --jsonexposes the completepaceobject. The TUI renders from the complete in-memory model, so--jsontiering never removes anything it draws. - When Codex reports banked rate-limit resets, the Codex card's weekly row ends with the count (
1 reset,3 resets), and the bar shortens to make room. A count above 99 shows as99+ resetsso the card keeps its width; TOON and JSON keep the exact count. A Codex card with no weekly window puts the count on its last row from the account-widerate_limitblock instead, never on a per-model or code-review row; only a card with no account-wide row puts it on its last row. When Codex does not report a count, or the card is stale or reused, the row is unchanged; no0 resetsis invented. - When Z.ai Coding Plan reports banked reset cards, each count sits at the end of its own window row (
4 resetson the session row,3 resetson the week row), and that row's bar shortens so the card keeps its width. A count above 99 shows as99+ resets. A stale or reused card, or one whose reading carries no counts, renders the rows unchanged; no0 resetsis invented. - A provider whose window relationships are wholly unknown (Copilot or Antigravity, with every window unresolved) has no combined effective percentage, pace, or runway to show, so its card replaces the headline block with a single
per-window usage · no combined boundline and leads straight into its real per-window rows. Partially understood providers keep the effective-unknown headline. No combined headroom, pace, or runway number is invented. - Providers are grouped by what quota-axi found on this machine, and the header always counts all four groups in that order, for example
4 live · 1 stale · 2 need attention · 10 not set up, including the ones that are empty. Every count is kept even in the narrowest supported terminal: a header that would not fit gives up its timestamp - the time zone first, then the date, then the clock - never a count. A fresh reading is live: an accented●and a full border. A stale cached reading is its own lane, counted apart from live. Its card uses a◐glyph and a dim border, and notes name when the snapshot was last refreshed and the error from the live read that failed, includingfetch failedwhen the failure was the usage fetch. A provider without a reading whose sources found something (a credential that is expired, rejected, or waiting on a Keychain prompt, an installed tool that failed, or a request that failed) needs attention: it stays a dimmed full card right after the live and stale cards, with its remedy, and is excluded from the live total. A provider is not set up only when every credential source it checks was skipped as absent (no file, environment variable, CLI, or running app). A provider that recorded no sources at all still needs attention, so absence is never assumed, and so does one whose skip could not establish absence either way, such as a Copilot CLI account whose Keychain value still awaits--allow-keychain-promptor a configuration that names no account quota-axi can confirm, a Kimi Code configuration quota-axi could not read to the environment it names, or an Antigravity CLI that is installed but timed out. A GitHub CLI login alone, or one Copilot refused, is not evidence of Copilot access and leaves Copilot not set up rather than broken; aghstore that could not be read, or a request through it that failed transiently, keeps Copilot in view. - Providers that are not set up fold into one dim footer line that names each of them and points at
quota-axi authfor where each one is read. Pressain the live report to draw them as dimmed full cards under a○ not set up · Nlabel, and press it again to fold them; while any are folded or expanded the closing line names the key, including when the report scrolls.--allstarts the report expanded and is how a--onceframe shows them. A provider named with--provideris never folded. - Width comes from the terminal, clamped to 80-120 columns; below the two-up width the grid reflows to one column. Color honors
NO_COLOR,TERM=dumb, and non-TTY stdout (the glyph skeleton is kept), re-enables withFORCE_COLOR, and uses truecolor whenCOLORTERMadvertises it, falling back to 256-color then ANSI-16. --tuicomposes with--providerscoping,--all, and--full(account identity and source-attempt footers, including for folded providers). It is mutually exclusive with--jsonand only supported by thequotacommand;--allis only supported with--tui.
Multiple accounts
A normal invocation reports every Codex ChatGPT subscription it can discover from sibling entries in one Pi auth.json.
It keeps each account's quota windows, resets, plan, effective availability, runway, and spendPriority separate.
Each account gets its own TUI card, naming its key on an account <key> line under the card title, including accounts whose quota cannot be read.
The default key an expanded report fills in for single-account providers is a schema artefact, so the TUI leaves it out of both the card and the --full footer; TOON and JSON still publish it.
quota-axi --provider codex --json
quota-axi --provider codex --tui --no-credential-refresh
quota-axi --provider codex --full --json # adds vendor identity and source attempts per keyPi's built-in provider id is openai-codex.
pi-codex-accounts stores additional logins under ids such as openai-codex-work in the same file.
quota-axi enrolls those already-present keys; it does not read codex-accounts.json, copy tokens, launch Pi, or change the active account.
Discovery order is the built-in openai-codex entry, then other openai-codex-* keys in lexical order.
Two keys that carry the same stored accountId are the same ChatGPT account and are not reported as extra capacity.
The later key stays a credential fallback until a probe succeeds or every candidate is rejected.
The lane keeps the first key as its accountKey, while source names the key that answered.
See Account keys and compatibility for how consumers join folded keys to a row.
A key whose identity cannot be compared is left as its own lane so the uncertainty stays visible.
When only the built-in Pi entry (or none) is present, Codex keeps its existing single-winner path: native $CODEX_HOME/auth.json, then openai-codex, then the CLI fallback.
That row's accountKeys names the key of the credential that produced it: codex-home for the native login or CLI fallback, openai-codex for the built-in entry. A failed row names the credential it speaks for, and a stale row names the one that produced its cached snapshot. The other key is added when the native login and the built-in entry store the same accountId. A --profile-only row is always codex-home.
When siblings are present and a native $CODEX_HOME/auth.json exists, that login stays first as its own codex-home lane, read from auth.json and then the CLI fallback.
The built-in openai-codex entry whose stored accountId matches the native login is not a separate lane; it stays the native lane's fallback, as it was before.
Without a native auth.json, an installed Codex CLI fallback is probed once as the codex-home lane.
The lane is left out only when the app-server's account/read positively reports no ChatGPT login (account: null or a non-ChatGPT account).
A reading without the optional accountId stays its own lane with no identity.
A failed CLI reading is shown as stale or unavailable only when account/read confirmed a ChatGPT login or a codex-home snapshot is cached; a probe that fails before that evidence adds no lane, and auth still shows the cli-rpc source.
A native login (from auth.json or the CLI) for the same account as a Pi lane is not a second lane.
The account is compared by the vendor accountId a fresh reading reports, or else the stored one; email, tokens, and key names are never used as identity.
That Pi lane keeps its own reading when fresh, and shows the native reading when its own is expired, rejected, or stale and the native one is fresh, or when only the native one has stale cached windows.
A proven sign-out retires only cached Codex snapshots whose stored account identity matches the rejected credential; a native login that coalesces into a Pi lane no longer removes the shared codex-home slot by name, so another account's snapshot cannot be discarded.
--profile-only still reads one native Codex file and never opens Pi auth.
Account keys and compatibility
When a provider expands to multiple accounts, the report uses quota schemaVersion: 6 (auth and models use version 2).
Every provider record then has an accountKey; providers still using one selected account use the literal default.
Every quota-axi output quota row, in schema 5 and schema 6 alike, has accountKeys: the credential keys that one row covers, own accountKey first, then any keys folded into it in the order the provider recorded them. The exported ProviderQuota.accountKeys type remains optional so package consumers can construct reports without it; quota-axi output sets it on every quota row.
A row covering one credential lists just its own key. A consumer that holds a credential key binds the row whose accountKeys contains that key. Matching accountKey alone misses a key that was folded into another row.
A provider without account discovery lists accountKeys: ["default"], the same literal as its schema 6 accountKey filler. A discovering provider that stays on one row, such as two Pi keys for one account with no native login, keeps schema 5 and no accountKey, and its accountKeys starts with that lane's key.
Membership mirrors quota-axi's own grouping exactly: every key the provider grouped into the row, and no other. That grouping relies on stored account identity when no fresh live reading confirms it, so a fold is published as applied even when a later live reading names a different account.
accountKeys is a quota JSON field. Auth still lists each discovered lane on its own, because a fold that depends on the quota reading has not happened there. TOON does not add a column: its flat blocks already name the published lane by accountKey, and the membership list is the JSON account row's join field.
The shared account collector puts each discovered lane's key first; an adapter adds folded keys to its quota reading's ProviderQuota.accountKeys, including keys learned during the read. Codex is the only adapter that discovers accounts today; other providers report default. The list itself is not cached; Codex reconstructs a stale row's originating key from the cached credential source.
Every flat TOON block adds accountKey immediately after provider, and the quota/exhaustion/attention join becomes provider + accountKey + scope.
Models and model sort ties use provider + accountKey + id.
Models unmatchedWindowIds entries gain the same key, so an unmapped window reads provider/accountKey/scope instead of provider/scope; the key keeps two accounts of one provider from reporting the same unmapped window indistinguishably.
Declaration order remains non-preferential; quotas are never combined across accounts.
A Codex Pi lane's key is the auth.json provider id (openai-codex, openai-codex-work); the native Codex lane's key is codex-home.
It is stable across refreshes and discovery order and contains no token, email, or path, and it also names the account's cache slot.
A lone lane keeps the legacy keyless slot, which the single selected account uses too, so the snapshot itself records the stored ChatGPT account id of the credential that produced it (see Cache).
A key the report cannot publish (malformed or repeated) costs only its own lane: the lanes with usable keys still expand, so one unreadable entry never hides the accounts beside it.
The provider falls back to its single selected account only when no usable lane remains.
--full adds the vendor identity the usage endpoint supplied, when any.
If no provider expands, existing fields retain their shape and order, while quota JSON adds accountKeys: quota schema 5, auth/models schema 1, and no account column.
A sole discovered Pi sibling uses that legacy representation.
Expansion follows the lanes discovered rather than the rows published, so when a native login and a Pi sibling turn out to be one account the single surviving row still carries its key and the report stays schema 6.
Consumers must honor the schema version; a legacy keyless row means the single selected lane, and keys must never be inferred from row position.
Account collection is shared in src/providers/accounts.ts.
Adapters can implement ProviderAdapter.discoverAccounts with stable keys and bound quota/auth readers; collection preserves each account's success or failure.
Codex Pi sibling entries are the first discovery implementation on this tree.
Other adapters retain their existing source-selection behavior.
Output Model
The quota command's --json emits schemaVersion: 5, or 6 when a provider expands to multiple accounts.
Normalized schema contract
The package publishes TypeScript declarations from its package root, so consumers can use import type { QuotaAxiResponse, ModelsResponse } from "@zaphodis42/quota-axi". The adapter contract is ProviderAdapter in and normalized ProviderQuota out: adapters report observed quota data, never rank, mint credentials, or retain raw responses. The narrowly bounded vendor-owned renewal path is documented under Delegated credential refresh.
schemaVersion is command-specific. Additive fields do not bump it. A semantic or incompatible shape change does. The legacy single-account quota report is version 5, auth is version 1, and models is version 1. When account discovery expands a provider, those versions are 6, 2, and 2 respectively. For quota row membership, see Account keys and compatibility.
Default report blocks
Default TOON is organized by the reading agent's decision rather than by quota-axi's data structures:
| Block | Rows |
| -------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| quota[] | One row per measurable scope: provider, optional accountKey, scope, effectivePercentRemaining, spendPriority, runway, confidence, limitedBy, resetsAt. Every column is populated on every row. limitedBy is the scope's limitingWindowIds, and resetsAt is that binding window's own reset. |
| exhaustion[] | Sparse. One row per scope with a finite exhaustion point: usableRunwaySeconds, projectedExhaustedAt, limitingWindowId. exhaustion[0]: means nothing is projected to run out. |
| attention[] | Sparse. Every non-nominal fact: provider, optional accountKey, scope, kind, detail, remedy. |
A quota[] row whose runway is projected_exhaustion or exhausted_now has exactly one matching exhaustion[] row, joined on provider + scope (plus accountKey in an account-expanded report). A row with through_reset or unknown has none, by definition: through_reset deliberately has no deadline and unknown has none to state.
attention[] kinds:
| kind | scope | Meaning |
| ------------------------------------------------------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| stale | all | The report is stale diagnostic data. detail names the last refresh, fetch failed plus state.error when a usage fetch failed, and any state.reason; no scope gets a quota[] row. |
| auth_required, rate_limited, unavailable, error | all | The provider state status. detail is state.error, any state.reason, plus the retry-after instant when state.retryAfter is set. |
| no_quota | all | The provider reported no measurable scope and no raw credit balance. Emitted when nothing else names it or when needed to preserve state.authStatus. |
| credits | all | The provider reported a raw credit balance but no measurable scope. detail states that balance verbatim; no percentage or bound is derived from it. |
| resets_available | all | The provider reported banked rate-limit resets still available: Codex's rate_limit_reset_credits.available_count, or Z.ai Coding Plan's banked reset cards counted per reset type. detail states the vendor counts verbatim (plus the earliest Z.ai card expiry). Only a fresh reading reports it: the row is omitted when the vendor sends nothing, and on a stale or reused reading, because a banked reset may have been spent or granted since. The counts are never cached. |
| unresolved_windows | all | quotaSemantics.unresolvedWindowIds: unfamiliar vendor windows not folded into any bound. |
| untrusted_windows | all | state.untrustedWindowIds: limits that could not be parsed authoritatively. |
| share | all | A window is a used-share of another window, not an independent allowance. detail is <id> of <parent> plus · <percentUsed> when that figure is present. It bounds no scope. |
| headroom_unknown | scope | The scope reports no effective percentage for a reason other than a bound conflict. detail names the windows that block it and any finite runway verdict with its limiting window. |
| bound_conflict | scope | A window the scope only inherits reads zero while the scope's own windows still report allowance. detail names both sides. The scope gets no quota[] row and no exhaustion[] row. |
| unmeasurable | scope | Headroom is known but a bound blocks runway, spendPriority, or both. detail names which. |
| degraded_source | all | A credential source was superseded: it was broken or unreadable while a sibling source answered. detail is <source> · <error>. One row per source, only on a fresh reading. |
| reused | all | A fresh reading was reused instead of asking the vendor again. detail is last refreshed <refreshedAt>. See Fresh reuse. |
remedy carries state.remedyCommand when one exists, and situational agent-directed advice is still prepended to help.
Two invariants hold for every report:
- Every requested provider appears at least once, in
quota[]orattention[], except providers with positive evidence of absence, which default TOON omits and counts in a help line;--fulland--providername them. A provider that remains in the report with noquota[]row always states itsstate.authStatus- including a positiveusable- as(auth <status>)in itsattention[]detail. In an account-expanded report a provider is omitted only when every one of its lanes is absent. quota[]rows stay in provider-declaration order, never sorted by any metric. A compact table with aspendPrioritycolumn must never read as a published ranking.
The omission help line sits after any situational advice and before the tier hint:
10 providers not set up are omitted; run `quota-axi --full` to list themOne omitted provider uses the singular (1 provider, is, it). A machine with nothing set up still prints quota[0], exhaustion[0], and attention[0] plus that line, and still exits 1. --full prints the omitted rows and no omission line. Exit codes are otherwise unchanged.
An unknown or stale scope deliberately gets no quota[] row: the absence of a number is the correct encoding of "no number", and the scope is named in attention[] instead.
Output tiers
--full adds; it never subtracts. Default TOON carries the three decision blocks; --full TOON adds the providers[], windows[], scopeAudit[], accounts[], and attempts[] audit blocks. Default --json carries the normalized model with derivation inputs demoted; --full restores them with no renames and no re-nesting - a demoted field is simply absent until --full, in the exact position and under the exact name it has there.
| Demoted to --full in --json |
| ---------------------------------------------------------------------------------------------------------------------------------------------- |
| providers[].label, providers[].source |
| state.refreshedAt, state.sourcesTried |
| windows[].percentUsed (kept when the window carries shareOf, where it is the only figure), windows[].startsAt, windows[].windowSeconds |
| windows[].pace.timeRemainingPercent, elapsedPercent, cycleBasis, cycleSeconds, projectedExhaustedAt, projectionConfidence |
| quotaSemantics.description |
| effectiveAvailability[].pace.behindWindowIds, onPaceWindowIds |
| Account identity (account) and per-source attempts
