pi-ask-side
v1.0.0
Published
Ask the LLM token-efficient side questions in an isolated context while the main agent keeps working. Three modes: no context, last exchange (+ on-demand fetch), and deep (full session context + read-only repo tools).
Maintainers
Readme
ask-side
Ask the LLM a side question in an isolated, token-efficient context while the main agent keeps working. Three modes:
- No context (default) — the question goes to the model alone: no session
history, no tools, a tiny system prompt, thinking off, and
cacheRetention: "none"so it never touches the main session's prompt cache. - Last exchange (+ on-demand fetch) — the question is answered with the
most recent finalized user/assistant exchange from the main session as
context (text only; tool I/O is skipped). If the model says it needs more
prior context, the extension fetches one more exchange and asks again
(capped at 6 rounds / ~8k tokens of history). You can also press
mto force one more fetch round. Mode 2 reuses onesessionId+ short prompt caching so the accumulated prefix is cached across rounds. - Deep (full context + repo tools) — for when modes 1–2 can't answer from
the tail of the conversation. Runs a private agent loop seeded with
exactly the messages the main agent would send (compaction- and
branch-aware, tool calls/results included) plus pi's real system prompt
(incl. AGENTS.md), and gives it repo tools — just like a normal chat
session, but fully isolated. Calls go through the live model registry,
so custom providers registered by extensions (tokenhub, clinepass, …)
work. Tool activity streams live into the panel. The side conversation
stays alive while the panel is open, so
mfollow-ups are cheap incremental turns (stablesessionId+ short cache → prefix is a cache read). Capped at 12 turns per question.
Nothing a side question produces is added to the main agent's context unless
you press s to promote it into a steering message. The exchange is logged in
the transcript as a TUI-only entry (the model never sees it).
Trigger
| Trigger | Mode | Notes |
|---|---|---|
| Alt+A | 1 (no context) | Works while the agent is mid-stream. |
| Alt+W | 2 (last exchange) | Works while the agent is mid-stream. |
| Alt+E | 3 (deep: full context + tools) | Works while the agent is mid-stream. |
| Alt+M | (in panel) cycle | Switch mode 1 → 2 → 3 → 1 before submitting. |
| /ask [text] | 1 | Direct if text given; else opens the panel. |
| /askc [text] | 2 | Direct if text given; else opens the panel. |
| /askd [text] | 3 | Direct if text given; else opens the panel. |
Alt+A/Alt+W/Alt+E also appear in /hotkeys. Panels always open in the
chosen mode (there's no "remember last mode" — Alt+A//ask is always mode
1, Alt+W//askc is always mode 2, Alt+E//askd is always mode 3).
Why
Alt+A?Alt+Qconflicts with pi's built-inapp.message.dequeue(restore queued messages to the editor) on Windows/WSL.Alt+Ais free in both the app- and editor-level default keymaps;Alt+W/Alt+E/Alt+Mhave no built-in conflicts.
Panel keys
Input phase
- type your question ·
Enterto ask ·Shift+Enternewline Alt+M(orAlt+A/Alt+W/Alt+E) switch mode ·Esccancel
Thinking phase
Escto abort the in-flight call- mode 3 shows live tool activity (
→ read: src/foo.ts…)
Answer phase
↑↓PgUpPgDnHomeEnd— scroll the answerm— mode 2: force one more fetch round (more prior context) · mode 3: ask a follow-up in the same deep session (cheap incremental turn; the seeded full history + previous side Q&A stay cached)s— send the Q&A to the main agent as a steering message (delivered after the current turn; or starts a turn if the agent is idle). The only path that adds tokens to the main context — opt-in.c— copy the answer to the clipboard (pbcopy/wl-copy/xclip)Esc/Enter— close (and log the exchange to the transcript as a TUI-only entry the LLM never sees)
How mode 3 (deep) works
- Seeding:
buildSessionContext(getEntries(), getLeafId())— the exact message list the main agent would send (compaction and branch summaries applied, tool calls/results included). A trailing dangling tool sequence (possible when the main agent is mid-turn) is dropped so the provider accepts the history. - System prompt: the session's real system prompt (
ctx.getSystemPrompt(), incl. AGENTS.md/skills) + a short "side-question mode" addendum. - Loop:
ctx.modelRegistry.complete()with the enabledAgentTools (built from pi's own tool factories, bound to the session cwd). Tool calls are executed directly — no permission gates — which is why the default tool set is read-only (read,grep,find,ls). SetPI_ASK_TOOLS=bash,read,edit,write,grep,find,lsif you want write access (a warning is shown in the panel when write-capable tools are enabled). - Isolation: own message list, own
sessionId(short cache retention), nothing written to the main session or disk. Usage/cost is tracked and shown per answer. - Caps: 12 LLM turns per question (
MAX_DEEP_TURNS); ~4000 output tokens per turn (PI_ASK_MAX_TOKENS).
How mode 2 picks context
- Source:
ctx.sessionManager.buildContextEntries()— the active branch with compaction applied. Text only is taken from finalizeduser/assistantmessages (tool calls, tool results, and images are skipped); each message is truncated (~1.5k tokens). If the session has been compacted, older fetches reach the compaction summary (the summarized past). - Round r uses the newest
2rmessages (≈ one user+assistant pair per round). - The model is instructed to answer, or reply
{"need_more":true}if more prior context is required. The extension then fetches one more round and re-asks. When no more history is available (or the cap is hit) but the model still asks for more, it's told to answer with whatever it has. - Caps: 6 rounds, ~8k tokens of history (
MAX_CONTEXT_CHARS), then it answers and the panel shows "cap reached".mis disabled once the cap is reached or history is exhausted.
Install
This directory is a self-contained pi package. Pick whichever channel you prefer:
# from npm (after it's published — shows up in the pi package gallery)
pi install npm:pi-ask-side
# from a local clone of this repo
pi install ~/pi-config/agent/extensions/ask-side
# try it once without installing
pi -e ~/pi-config/agent/extensions/ask-sidepi install writes to user settings (~/.pi/agent/settings.json); add -l to
install project-local instead. Remove with pi remove npm:pi-ask-side (or the
path you used).
Development: this same directory is symlinked at
~/.pi/agent/extensions/ask-side/ (via pi-config/install.sh), so pi
auto-discovers it — edit the source and run /reload to hot-reload. Don't
install it as a package and keep the auto-discovered symlink on the same
machine, or you'll load it twice.
Publishing a new version
cd ~/pi-config/agent/extensions/ask-side
npm version patch # or minor/major
npm publishThe pi-package keyword in package.json is what makes it appear in the
pi package gallery. pi's bundled packages
(@earendil-works/*, typebox) are declared as peerDependencies so they're
never bundled into the tarball — pi provides them at runtime.
Configuration (env vars, all optional)
| Variable | Default | Purpose |
|---|---|---|
| PI_ASK_MODEL | the session's active model | provider/modelId to use a cheaper/faster model. Falls back to the active model (with a warning) if not found or unauthenticated. |
| PI_ASK_MAX_TOKENS | 2000 (modes 1–2), 4000 (mode 3) | Max output tokens per side answer. |
| PI_ASK_SYSTEM_PROMPT | a tiny built-in prompt | Override the mode-1 system prompt. |
| PI_ASK_TOOLS | read,grep,find,ls | Comma-separated built-in tools for mode 3. Read-only by default; tools run without permission prompts, so adding bash/edit/write is on you. |
The default model is the session's active model, which is guaranteed to have
working auth (API key or subscription). For reasoning models using the
deepseek/openai thinking formats (e.g. GLM-5.2), thinking is left OFF for
side questions (no reasoning tokens). Set PI_ASK_MODEL=anthropic/claude-haiku-4-5
(or any fast/cheap model you have auth for) to make side questions even cheaper.
