pi-l1-cache
v1.4.0
Published
L1 response cache for pi with replay, disk persistence, and working capture
Maintainers
Readme
pi-l1-cache
L1 In-Memory Cache Extension for pi — with disk persistence, response replay, and working capture
A high-performance, production-ready L1 (in-memory + disk) cache extension for the pi coding agent. Unlike the stock API which lacks response capture, this implementation uses a pi-core patch to enable full caching with replay.
⚡ Performance
| Scenario | Cold | Warm (replayed) | Speedup | |---|---|---|---| | Simple prompt | 3.2s | 1.9s | ~1.7× | | Tool-calling agent loop | 13.1s | 2.5s | 5.2× |
Features
| Feature | Description |
|---------|-------------|
| Response replay | Cached chunks are fed through pi-ai's normal consume path — parsing, usage, tool-calls, stop-reason all identical |
| Disk persistence | Survives across pi -p process restarts (~/.pi/agent/cache/l1-cache/) |
| ~138µs overhead | End-to-end per-request cost (measured, incl. key hashing + JSON) |
| Fast key hash | FNV-1a (~1.5µs/op) — no per-request CPU probe on the hot path |
| Memory cap | Hard limit (20MB default) prevents RAM bloat |
| Auto-eviction | LRU-style cleanup when limits reached |
| TTL-based | 1-hour default expiry for cached entries |
| Coalescing | Adjacent content/reasoning_content deltas merged to shrink storage |
Architecture
pi → [L1: RAM Map + disk] → Provider
~138µs lookup 1-3sThe extension uses a pi-core patch (fix-l1-cache.cjs) to add two capabilities the stock extension API lacks:
Bundle era (pi ≥ 0.84): pi runs from the esbuild bundle (
dist/bundle/cli.js), sofix-l1-cache.cjspatchesdist/bundle/chunks/*(minified, anchor-matched — verified against 0.86.0). Pre-bundle pi (< 0.84) is patched in the readable pi-ai sources. The patch is idempotent and bundle-era misses degrade to pass-through rather than failing an install.
REPLAY — an extension may serve a cached response by returning params with
__piL1Replay: [chunks]frombefore_provider_request; pi-ai feeds the cached chunks through the normal consume path and never contacts the provider.CAPTURE — after a successful completion, pi-ai calls
options.onStreamComplete(allChunks, requestParams);sdk.jsforwards them to extensions as aprovider_stream_completeevent so the cache can store the response.
Installation
Option 1: From npm (recommended)
pi install npm:pi-l1-cacheThis installs the extension AND automatically applies the fix-l1-cache.cjs pi-core patch.
Option 2: From GitHub
pi install git:github.com:tobias-weiss-ai-xr/pi-l1-cache@mainOption 3: From local clone
pi install /path/to/pi-l1-cacheThen restart pi — the extension auto-loads.
Configuration
Environment Variables
# Enable/disable
L1_CACHE_ENABLED=true
# Max entries (default: 200)
L1_CACHE_MAX_ENTRIES=200
# Max memory in MB (default: 20)
L1_CACHE_MAX_MB=20
# TTL in seconds (default: 3600)
L1_CACHE_TTL=3600
# Log stats on each hit/miss (default: false)
L1_CACHE_LOG=trueSettings (in ~/.pi/settings.json)
{
"extensions": {
"l1-cache": {
"enabled": true,
"maxEntries": 200,
"maxMemoryBytes": 20971520,
"ttlSeconds": 3600,
"persist": true,
"logStats": false
}
}
}Usage
Show cache stats
/l1-cacheOutput:
L1 cache: 16 entries, 43.2KB | hits 12 (replays 12), misses 4, writes 4, evictions 0 | dir: C:/Users/Tobias/.pi/agent/cache/l1-cacheClear cache
/l1-cache clearCleared: memory + disk (all .json files in cache dir).
Key Semantics
The cache key is a stable hash of:
modelmessages(full conversation history)toolstool_choicetemperature,top_preasoning_effort,thinkingmax_completion_tokens,max_tokens
Volatile fields are EXCLUDED: prompt_cache_key, prompt_cache_retention, stream, stream_options, store, sessionId.
This means:
- ✅ Cross-run hits possible (same prompt, same cwd, same model)
- ✅ Tool-calling agent loops fully replayed (tool results are part of messages)
- ❌ Different conversation history = different key (expected)
- ❌ Session-derived fields don't break cross-run hits
Gotchas
Replays are canned — identical input returns the stored response verbatim. For fresh answers, use
/l1-cache clear.Print mode (
pi -p) reads entire stdin as ONE prompt — you cannot test two identical requests in one process this way.llm-timestamp.js does NOT pollute provider messages (it only appends display-level
message_endentries), so keys are stable across runs.Replayed responses carry original responseId/usage — accurate since input identical.
Troubleshooting
"fix-reasoning-content.js: layout changed"
The fix-l1-cache.cjs patch modifies the same file as fix-reasoning-content.js. The wrapper's postinstall chain handles this correctly — fix-reasoning-content.js now recognizes its work via marker even after the replay branch is added.
If you see this error:
- Ensure you're running the full postinstall chain (not individual scripts)
- Check that
fix-reasoning-content.jshas the marker-based detection (v1.2.3+) - Verify postinstall order:
fix-reasoning-content.jsBEFOREfix-l1-cache.cjs
Cache never hits
Check:
L1_CACHE_LOG=trueto see HIT/MISS/STORED logs- Prompt is byte-identical (including system prompt, tools, cwd context)
- Disk persistence is enabled (
persist: true) - TTL hasn't expired (default 1h)
Development
Testing
# Run unit tests
npm test
# Run the extension manually
npx tsx src/index.tsBenchmark
# Measure cold vs warm times
echo "What is 7*6?" | time pi -p "test" # cold
echo "What is 7*6?" | time pi -p "test" # warm (should be ~1.7× faster)License
MIT — see LICENSE
Contact
- Issues: https://github.com/tobias-weiss-ai-xr/pi-l1-cache/issues
- Email: [email protected]
