npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

pi-cache-metrics

v1.1.0

Published

Live prompt-cache visibility in the Pi footer — hit/miss rates, token usage, and estimated cost savings

Readme

Cache Metrics

Live prompt-cache visibility in the Pi footer. Tracks hit/miss rates, cache token usage, and estimated cost savings across the session.

What is prompt caching? Providers reuse your context prefix (system prompt + tools + conversation history) across requests. A cache hit (high hit-rate) means those tokens were served from cache instead of reprocessed — faster responses, lower cost. Misses occur when previously-cached tokens are re-billed as fresh input.

Install

Via npm:

pi install npm:pi-cache-metrics
# run directly: pi -e npm:pi-cache-metrics

Commands

| Command | Description | |---------|-------------| | /cache | Show session summary (hit rate, tokens, savings, per-model breakdown, trend chart) | | /cache-log | Show recent requests (last 10 of 50 kept) with input/read/write/saved, timestamps with ms | | /cache-settings | Open interactive TUI settings (up/down to navigate, Enter to cycle, Esc to close) |

/cache example output

╭─ Cache Metrics — session summary
│  Model: deepseek/deepseek-v4-flash-0731
│  Requests: 50  │  Hits: 41  │  Misses: 9  │  No-cache: 0
│  Hit rate: 82.0%
│  Cache read:  4.4M  │  Cache write: 0
│  Est. savings: $0.053
├─ Per-model:
│  deepseek/deepseek-v4-flash-0731: 50 req  │  82% hit  │  $0.053 saved  │  read 4.4M  │  write 0
├─ Settings:
│  Footer: on (compact)  │  Cost precision: 4  │  Alert: 50%
│  Track models: all  │  Color: auto  │  Chart: hit
╰────────────────────────────────────────────────────────────

╭─ Cache Metrics — session summary
│  Model: anthropic/claude-sonnet-4-5
│  Requests: 12  │  Hits: 12  │  Misses: 0  │  No-cache: 0
│  Hit rate: 100.0%
│  Cache read:  1.2M  │  Cache write: 450k
│  Est. savings: $0.018
├─ Per-model:
│  anthropic/claude-sonnet-4-5: 12 req  │  100% hit  │  $0.018 saved  │  read 1.2M  │  write 450k
├─ Settings:
│  Footer: on (compact)  │  Cost precision: 4  │  Alert: off
│  Track models: all  │  Color: auto  │  Chart: read
╰────────────────────────────────────────────────────────────

hit-rate trend (oldest → newest):
▁▁▂▃▄▅▆▇███▇▆▅▄▃▂▁▁▁▂▃▄▅▆▇███▇▆
  ▁ low · █ high · . no data

/cache-log example output (timestamps with ms, includes input tokens)

╭─ Cache Metrics — recent requests (last 10 of 50 kept)
│  14:32:15.123  deepseek-v4-flash  hit    in  98.6k  read  88.6k  write  0.0  $0.0021
│  14:32:18.456  deepseek-v4-flash  hit    in  98.6k  read  88.6k  write  0.0  $0.0021
│  14:32:45.789  deepseek-v4-flash  miss   in 102.4k  read   0.0   write  0.0  $0.0000
│  14:32:48.012  deepseek-v4-flash  hit    in 102.4k  read 100.0k  write  0.0  $0.0024
╰────────────────────────────────────────────────────────────────────────

/cache-settings (TUI)

The interactive settings panel includes a short on-screen explanation of caching:

| Setting | Values | |---------|--------| | Show Footer | on / off | | Footer Format | compact / detailed | | Cost Precision (decimals) | 0–6 | | Hit-Rate Alert Threshold (%) | off / 10–100 | | Color Theme | auto / mono / color | | Chart Metric | hit / read / saved |

  • hit (default) — hit rate % per time bucket
  • read — cache-read token volume per bucket (normalized to max)
  • saved — estimated savings per bucket (normalized to max)

Features

  • Provider-agnostic — uses Pi's normalized Usage (cacheRead, cacheWrite, cost.cacheRead, cost.cacheWrite), not fragile response headers
  • Robust miss detection (Pi-style) — compares consecutive requests: if tokens from the previous prompt were NOT read from cache this turn, it's a miss. Catches DeepSeek/OpenAI providers that bill cache misses as plain input instead of reporting cacheWrite.
  • Live footer — 💾 cache 67% (8H/4M) · cr12k/cw8.2k · -$0.0042
  • Color-coded — green ≥70%, yellow 40-69%, dim <40% (mono mode available)
  • Branch-safe persistence — state stored as appendEntry custom entries, survives /reload and session forks
  • Hit-rate alerts — optional notification when rate drops below threshold
  • Per-model breakdown — /cache shows aggregates per model
  • ASCII trend chart — switchable metric (hit / read / saved); blocks ▁→█ relative to max; . = no data
  • Bounded log — /cache-log keeps the last 50 requests in memory, shows the latest 10 with ms timestamps and input tokens

Data Source

Hook message_end → event.message.usage. Pi normalizes per-provider cache tokens (Anthropic cache_read_input_tokens, OpenAI cache usage, etc.) into a single Usage object. This is more robust than parsing provider-specific headers, which the docs note are not uniformly exposed across providers/transports.

Miss detection mirrors Pi's internal cache-stats.js: for each request, missedTokens = min(prev.promptTokens, promptTokens) - cacheRead. If missedTokens > 1024 and the previous request reported cache activity, it's a miss — even if cacheRead > 0 (partial miss). This catches providers like DeepSeek that never populate cacheWrite but re-bill the prompt as input on cache expiry.

Savings are estimated by looking up the model's input rate from ctx.model, the model registry, or inferring from usage — whichever is available.

Self-Test

CACHE_DEBUG_SELFTEST=1 pi -e ./cache-metrics.ts -p '1+1'
# [cache-metrics] selfTest: 21 assertions passed

Or standalone (no Pi, no keys):

bun run cache-metrics.test.ts

License

MIT