pi-cache-metrics
v1.1.0
Published
Live prompt-cache visibility in the Pi footer — hit/miss rates, token usage, and estimated cost savings
Maintainers
Readme
Cache Metrics
Live prompt-cache visibility in the Pi footer. Tracks hit/miss rates, cache token usage, and estimated cost savings across the session.
What is prompt caching? Providers reuse your context prefix (system prompt + tools + conversation history) across requests. A cache hit (high hit-rate) means those tokens were served from cache instead of reprocessed — faster responses, lower cost. Misses occur when previously-cached tokens are re-billed as fresh input.
Install
Via npm:
pi install npm:pi-cache-metrics
# run directly: pi -e npm:pi-cache-metricsCommands
| Command | Description |
|---------|-------------|
| /cache | Show session summary (hit rate, tokens, savings, per-model breakdown, trend chart) |
| /cache-log | Show recent requests (last 10 of 50 kept) with input/read/write/saved, timestamps with ms |
| /cache-settings | Open interactive TUI settings (up/down to navigate, Enter to cycle, Esc to close) |
/cache example output
╭─ Cache Metrics — session summary
│ Model: deepseek/deepseek-v4-flash-0731
│ Requests: 50 │ Hits: 41 │ Misses: 9 │ No-cache: 0
│ Hit rate: 82.0%
│ Cache read: 4.4M │ Cache write: 0
│ Est. savings: $0.053
├─ Per-model:
│ deepseek/deepseek-v4-flash-0731: 50 req │ 82% hit │ $0.053 saved │ read 4.4M │ write 0
├─ Settings:
│ Footer: on (compact) │ Cost precision: 4 │ Alert: 50%
│ Track models: all │ Color: auto │ Chart: hit
╰────────────────────────────────────────────────────────────
╭─ Cache Metrics — session summary
│ Model: anthropic/claude-sonnet-4-5
│ Requests: 12 │ Hits: 12 │ Misses: 0 │ No-cache: 0
│ Hit rate: 100.0%
│ Cache read: 1.2M │ Cache write: 450k
│ Est. savings: $0.018
├─ Per-model:
│ anthropic/claude-sonnet-4-5: 12 req │ 100% hit │ $0.018 saved │ read 1.2M │ write 450k
├─ Settings:
│ Footer: on (compact) │ Cost precision: 4 │ Alert: off
│ Track models: all │ Color: auto │ Chart: read
╰────────────────────────────────────────────────────────────
hit-rate trend (oldest → newest):
▁▁▂▃▄▅▆▇███▇▆▅▄▃▂▁▁▁▂▃▄▅▆▇███▇▆
▁ low · █ high · . no data/cache-log example output (timestamps with ms, includes input tokens)
╭─ Cache Metrics — recent requests (last 10 of 50 kept)
│ 14:32:15.123 deepseek-v4-flash hit in 98.6k read 88.6k write 0.0 $0.0021
│ 14:32:18.456 deepseek-v4-flash hit in 98.6k read 88.6k write 0.0 $0.0021
│ 14:32:45.789 deepseek-v4-flash miss in 102.4k read 0.0 write 0.0 $0.0000
│ 14:32:48.012 deepseek-v4-flash hit in 102.4k read 100.0k write 0.0 $0.0024
╰────────────────────────────────────────────────────────────────────────/cache-settings (TUI)
The interactive settings panel includes a short on-screen explanation of caching:
| Setting | Values |
|---------|--------|
| Show Footer | on / off |
| Footer Format | compact / detailed |
| Cost Precision (decimals) | 0–6 |
| Hit-Rate Alert Threshold (%) | off / 10–100 |
| Color Theme | auto / mono / color |
| Chart Metric | hit / read / saved |
hit(default) — hit rate % per time bucketread— cache-read token volume per bucket (normalized to max)saved— estimated savings per bucket (normalized to max)
Features
- Provider-agnostic — uses Pi's normalized
Usage(cacheRead,cacheWrite,cost.cacheRead,cost.cacheWrite), not fragile response headers - Robust miss detection (Pi-style) — compares consecutive requests: if tokens from the previous prompt were NOT read from cache this turn, it's a miss. Catches DeepSeek/OpenAI providers that bill cache misses as plain input instead of reporting
cacheWrite. - Live footer —
💾 cache 67% (8H/4M) · cr12k/cw8.2k · -$0.0042 - Color-coded — green ≥70%, yellow 40-69%, dim <40% (mono mode available)
- Branch-safe persistence — state stored as
appendEntrycustom entries, survives/reloadand session forks - Hit-rate alerts — optional notification when rate drops below threshold
- Per-model breakdown —
/cacheshows aggregates per model - ASCII trend chart — switchable metric (
hit/read/saved); blocks ▁→█ relative to max;.= no data - Bounded log —
/cache-logkeeps the last 50 requests in memory, shows the latest 10 with ms timestamps and input tokens
Data Source
Hook message_end → event.message.usage. Pi normalizes per-provider cache tokens (Anthropic cache_read_input_tokens, OpenAI cache usage, etc.) into a single Usage object. This is more robust than parsing provider-specific headers, which the docs note are not uniformly exposed across providers/transports.
Miss detection mirrors Pi's internal cache-stats.js: for each request, missedTokens = min(prev.promptTokens, promptTokens) - cacheRead. If missedTokens > 1024 and the previous request reported cache activity, it's a miss — even if cacheRead > 0 (partial miss). This catches providers like DeepSeek that never populate cacheWrite but re-bill the prompt as input on cache expiry.
Savings are estimated by looking up the model's input rate from ctx.model, the model registry, or inferring from usage — whichever is available.
Self-Test
CACHE_DEBUG_SELFTEST=1 pi -e ./cache-metrics.ts -p '1+1'
# [cache-metrics] selfTest: 21 assertions passedOr standalone (no Pi, no keys):
bun run cache-metrics.test.tsLicense
MIT
