pi-permission-classifier
v0.5.2
Published
Auto-classifier Authorizer chain link for pi-permission-system: a light model reviews each ask and returns allow, deny, or defer
Maintainers
Readme
pi-permission-classifier
An auto-classifier for the pi coding agent's
permission system. It registers an Authorizer chain link named classifier
with @gotgenes/pi-permission-system:
a light model reviews each permission
ask and returns allow, deny, or defer, so clearly benign requests are
approved automatically, clearly bad ones are rejected with a short teaching
reason, and only genuinely uncertain ones reach you.
Think of it as a Claude-Code-style auto-approve mode, without giving up the permission gate: your deterministic policy still runs first, and the classifier only sees the asks that policy would have sent to you anyway.
Fail-safe by construction: a missing config, an unresolved model, an auth failure, a timeout, an unparseable reply, or any internal error resolves to defer. More prompting, never less.
Requirements
- pi with
@gotgenes/pi-permission-system27.0.0 or newer installed and active in the same session (the classifier resolves the permission service through the session-keyed locator introduced in 27.0.0; with an older version it warns once and registers nothing) - pi 0.84.3 or newer for the searchable
/permission-modelpicker, declared as the@earendil-works/pi-coding-agent >=0.84.3peer dependency: the selector component that picker mounts changed constructor shape in 0.84.3, so on an older pi the command errors before the picker mounts - nothing is written, the judge keeps reviewing, and the typed andsessionforms still work. Everything else needs only the permission-system floor above. - Node 22 or newer
- No build step: pi loads
src/index.tsdirectly
Setup / quickstart
Everything below is an operator action - the package never enables itself, and installing it grants it no authority until you name it in the chain.
Install the package from npm:
pi install npm:pi-permission-classifierpi installs it under
~/.pi/agent/npm/and addsnpm:pi-permission-classifierto thepackageslist in~/.pi/agent/settings.json.Check the package order.
pi installadds the new entry at the end ofpackages, so it can land afterpi-permission-system. Edit~/.pi/agent/settings.jsonif needed. Listpi-permission-classifierbeforepi-permission-systeminpackages:"packages": [ "npm:pi-permission-classifier", "npm:@gotgenes/pi-permission-system" ]Order matters: pi runs
session_starthandlers in package order, andpi-permission-systememits itspermissions:readyevent from its ownsession_start. With the classifier listed first, the link and thezz-permission-classifierfooter entry are in place at startup. Listed after it, the classifier misses that first ready event and registers at the next one, so the link and the footer entry appear only after the first agent turn (asks are still reviewed, since they happen inside agent turns).To run from a local checkout instead, clone the repository, run
npm installin it, and put its directory inpackages(path relative to~/.pi/agent, or absolute), in the same position.Activate the chain link. In
~/.pi/agent/extensions/pi-permission-system/config.json, add:"authorizerChain": ["classifier"]Only links named here are consulted; config order fixes the chain order.
Create the classifier config. Without it the link registers nothing (a safe no-op). The defaults are a good starting point:
mkdir -p ~/.pi/agent/extensions/pi-permission-classifier echo '{}' > ~/.pi/agent/extensions/pi-permission-classifier/config.json{}means: judge with the session's active model, judge every surface exceptpathandexternal_directory, 5000 ms timeout, built-in rubric. Seeconfig/config.example.jsonfor a version with a dedicated judge model, or pick one later from inside pi with/permission-model(see "Choosing the judge model").Try it. Start a new pi session and trigger something your policy sends to
ask(for example a bash command not on your allowlist). A benign command should now be approved automatically; a command matching the never-allow list should be rejected with a reason; anything uncertain still prompts you.Watch the decisions. Every reviewed ask writes one
classifier.decisionentry to the permission review log:~/.pi/agent/extensions/pi-permission-system/logs/pi-permission-system-permission-review.jsonlEach entry records request id, surface, value, model id, latency, verdict, defer reason, and three context fields:
contextIncluded(true only when a full command was rendered to the judge),contextBytes(its UTF-8 byte length, null when absent), andcontextHash(first 12 hex chars of its sha256, null when absent). The full-command text itself is never logged. An over-budget context defers with reasoncontext-over-budgetbefore any model call. EnabledebugLogin the permission system config to also capture raw model replies and short-circuit traces.
To disable, remove classifier from authorizerChain (or remove the
package entry). The previous prompting behavior returns immediately.
Configuration
Config files (project overrides global, shallow merge):
- Global:
~/.pi/agent/extensions/pi-permission-classifier/config.json - Project:
<cwd>/.pi/extensions/pi-permission-classifier/config.json
| Field | Default | Meaning |
| --- | --- | --- |
| provider | unset | Judge model provider. Set together with model; with neither set, the session's active model judges. |
| model | unset | Judge model id, resolved from the session model registry. Set together with provider. |
| instructions | built-in rubric | System prompt for the judge. Replaces the default rubric verbatim when set. |
| surfaces | ignored | Accepted so older config files still parse, but ignored: every surface outside the path and external_directory families is judged. A file that sets it logs one warning at load: This field is ignored: every surface outside the path and external_directory families is judged. |
| timeoutMs | 5000 | Per-review model call budget in milliseconds (positive integer). |
| contextBudgetBytes | 8192 | Cap on the extracted full-command context in UTF-8 bytes (positive integer). An ask whose context exceeds the budget defers before any model call; context is never truncated to fit. |
A malformed or invalid config file means the link registers nothing and pi logs a warning - the gate falls back to normal prompting.
Example: a local judge model
A small local model makes a good judge: verdicts stay on your machine, cost
nothing, and return fast. Register the model in pi's models.json under a
local OpenAI-compatible provider (llama.cpp, llama-swap, Ollama, vLLM):
{
"providers": {
"llama-cpp-local": {
"baseUrl": "http://localhost:8080/v1",
"api": "openai-completions",
"apiKey": "no-key",
"models": [
{
"id": "gemma-judge",
"name": "Gemma 4 E4B (judge)",
"reasoning": false,
"contextWindow": 16384,
"maxTokens": 4096,
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 }
}
]
}
}
}Then point the classifier at it:
{
"provider": "llama-cpp-local",
"model": "gemma-judge"
}The model must support tool calling: the classifier forces a
report_verdict tool call and treats a reply without one as defer. Verify
with a quick curl ("tool_choice": "required") that your server returns
finish_reason: "tool_calls". Keep the model resident if you can - a cold
load on the first ask can eat the timeoutMs budget.
The default rubric
Balanced and defer-first: allow clearly benign, intent-aligned asks; deny only the hard never-allow list; defer anything uncertain. The never-allow list covers secret/credential access, exfiltration, pipe-to-shell installs, force push, discarding uncommitted work, disarming safety guards, and edits to the permission system's or the classifier's own config and logs.
Two refinements come from field use. Scripts the judge cannot see always
defer: a bare python3 or bash /tmp/run.sh reaches the judge as an
opaque interpreter call. Inline code (node -e, python3 -c) is visible
and judged on its content. And a plain git stash (no drop or clear)
is not treated as discarding work.
When the ask carries the enclosing full command (a gated unit inside a larger command line), the judge reads all of it and allows only when every part is clearly benign. An interpreter body visible in the full command - a heredoc body or text piped to stdin - counts as inline code and is judged on its content; an interpreter run on a script file stays unseen and defers.
Two more lines cover the asks that guidance files most often need to
raise. Network fetches: downloading from any host, localhost included, is
allow when the fetched bytes are only written to files inside the project
tree or /tmp and nothing executes them; a fetch that feeds a shell or
interpreter is the pipe-to-shell never-allow item, and any other
destination or use defers. Cleanup deletes: a plain rm of named files or
build output inside the project tree or /tmp is cleanup, not discarding
work; a delete that reaches outside those places, removes tracked changes,
or uses paths the judge cannot resolve defers.
The rubric adds one line on guidance files (next section): they describe what is normal for this operator and project, can move a verdict toward allow or deny within the rubric, and never override the never-allow list.
Set instructions to replace the rubric wholesale with your own.
Guidance files
The judge reads the same context files the agent does: the global
AGENTS.md (or CLAUDE.md fallback) in the pi agent dir, and the project
files pi finds from the session cwd up through its ancestors (override
names and worktree shadowing included, because the classifier calls pi's
own context-file loader). Write the lines that describe what is normal
here, for example "downloads into ./vendor and cleanup of build/ are
routine in this repo", and the judge sees them verbatim.
- Files are read from disk on every judged ask. An edit takes effect on the next ask; there is no cache and no reload command.
- Trust gating: the global file always reaches the judge. Project files
reach it only while pi reports the project trusted, read at the moment
of each ask. In an untrusted directory the judge sees the global file
alone, and the classifier's own project config layer
(
.pi/extensions/pi-permission-classifier/config.json) is not read either, so a projectinstructionsstring cannot replace the rubric before you trust the directory. - Caps: a file over 16 KiB (UTF-8 bytes) is dropped whole, never truncated. Files accumulate in loader order (global first, then root down to cwd), and once the running total would pass 32 KiB that file and every later one are dropped whole.
- Rendering: each included file appears in its own delimited data block
after the ask facts, labelled
Operator guidance from <path>for the global file orProject guidance from <path>for a project file, under a one-sentence header on what guidance may and may not do. Guidance is data to the judge, not instructions. The header and the rubric tell the judge the seven never-allow items are outside its reach, but that is a prompt instruction, not code: a trusted project's file can still steer a model-based verdict, so review what you trust. The engine's ownpathandexternal_directorycaps hold regardless. - Logging: every
classifier.decisionentry carriesguidanceIncluded(one{path, bytes, hash12}per rendered file) andguidanceDropped(one{path, bytes, reason}per excluded file, reason one ofover-file-cap,over-total-cap,untrusted). File content is never logged. - Failure: if the loader throws, the ask defers with reason
guidance-load-failed, recorded on the decision entry and shown in the footer health suffix like any other failure defer.
Config suggestions
The judge decides only what the pi-permission-system policy sends to
ask. A few policy choices, learned from the review log, keep the judge
useful and cheap:
- Allow the commands you run all day with static rules instead of a judge
call each time:
npm run *,npm test*,npx vitest*,npx tsc*. In one logged day the judge allowed these 30-plus times at 2-5 s each. Do not allownpx *broadly; it downloads and runs packages. - Remember that
*crosses/in bash patterns.rm -rf /*denies every absolute-pathrm -rf, including/tmp/scratch; write the exactrm -rf /instead and let the judge see the rest. The same applies torm -rf ~/*. - Send
find *-exec*toaskrather thandeny. The judge receives the executed unit (cat {}) and allows read-only uses. - Turn reasoning off on a local judge (llama-server
--reasoning-budget 0, or"enable_thinking": falsein the chat template kwargs). A thinking block of 250 tokens costs 2-5 s on a small GPU and hits the 5000 ms budget on long asks; without it a verdict returns in under 1 s with the same verdicts on the same asks. - Read the decision trail. The permission system's review log
(
logs/pi-permission-system-permission-review.jsonl) records oneclassifier.decisionentry per reviewed ask with the verdict, defer reason, and latency. A run ofdeferwith reasontimeoutmeans the model, not the rubric, needs attention.
Which surfaces are judged
The classifier judges every surface except the path and
external_directory families, whatever the surface name, including
surfaces added by other extensions. A family is the bare name plus its
<name>_* members, today path, path_read, path_write,
external_directory, external_directory_read, and
external_directory_write. There is no surface list to maintain. Your
cross-cutting path and external_directory rules still apply, and the
engine downgrades any link allow on those families to defer, so the
classifier can never approve access outside the working directory or to a
path your policy denies. To keep a tool out of the judge's hands, route it
to allow or deny in the permission system policy instead of ask.
Choosing the judge model
The judge is the model that reviews each ask. With no provider and
model in the config it is the session's active model. Three ways to
change it:
The /permission-model command
/permission-modelwith no argument opens pi's own searchable model picker (the same list as/model, scoped models included) with the current judge preselected. Pick a model to make it the judge; cancel to change nothing. Outside the TUI (rpc, json, print modes) the command prints the current judge and the usage line instead. The searchable picker needs pi 0.84.3 or newer (@earendil-works/pi-coding-agent >=0.84.3): its constructor changed in 0.84.3, so on an older pi the command errors before the picker mounts, writes nothing, and leaves the judge on its previous rule - loud and safe, not a silent fallback. If your pi version does not expose the registry runtime the picker needs, the command warns that the picker degraded and offers a plain list ofprovider/idlabels./permission-model <provider>/<id>sets the judge by reference. The pair must be in pi's model registry: an unknown pair is rejected and nothing changes. A known model without configured auth is accepted with a warning, and asks defer until the auth exists. Tab completion offers theprovider/idlabels of the available models plussession./permission-model sessionremovesproviderandmodelso the session's active model judges again.
A choice applies to the next reviewed ask immediately and is saved to the
global config file
(~/.pi/agent/extensions/pi-permission-classifier/config.json): only
provider and model are rewritten, every other field is preserved, and
the write goes through a temporary file that is fsynced and then renamed, so
neither a process crash nor an OS crash leaves a truncated config. The command
never creates the file. Every write form refuses - nothing is written, and the
setup hint names the global path - when the global config file is absent or
there is no valid merged config, which means no config file was found this
session or the files found failed validation. The precondition is the config,
not registration: with a valid config and a link that never registers
(pi-permission-system absent or older than 27.0.0) a write still succeeds,
and the new judge applies as soon as the link does register.
When the project config sets provider or model, the global write still
happens but the command warns that the project file shadows the choice in
that project.
Choosing a judge never changes pi's session model or your default model:
/model and /permission-model are independent.
The --permission-model launch flag
pi --permission-model <provider>/<id> makes that model the judge for the
session only. It takes precedence over the config and the session model,
nothing is written, and it is dropped at session shutdown. The flag is read
again at every session start in that pi process, so after /reload or a new
session in the same process it applies again. A reference the registry does
not know is ignored with a warning and the configured judge applies. An
explicit /permission-model choice during the session replaces the flag for
the rest of that session.
The status bar entry
Once the link registers, the footer shows the effective judge under the
key zz-permission-classifier, in one of three states (pi sorts extension
statuses by key on one footer line; the zz- prefix keeps the judge entry
last, at the end of that line):
judge:<provider>/<id>- a configured or flag-set judgejudge:session- the session's active model judgesjudge:<provider>/<id> (unresolved)- the configured pair is not in the registry, so every ask defers until it resolves
The entry follows /permission-model changes and /model switches and is
cleared at session shutdown. No entry means the link did not register.
The judge text carries a health suffix once something went wrong this session. It refreshes after every decision, so the footer is the quickest read on how often the dialog fell back and why:
- no suffix - no failure defers yet and the breaker is closed
| <reason> x<N>- the last decision was a failure defer with that reason (timeout,call-failed,context-over-budget,breaker-open,model-unresolved,auth-failed,no-tool-call,unrecognized-verdict,internal-error); N is the session's failure defer count| defers x<N>- a later model verdict cleared the reason; the count stays for the session| breaker open <S>s- the circuit breaker is cooling down, S is the remaining whole seconds, counted down once per second; this state wins over the others while it lasts
Model verdicts are not health events: an allow clears the pending reason, a deny changes nothing, and neither shows in the footer. The judge's own defer verdict counts as a model verdict. Everything resets at session shutdown.
How it works
- The classifier only sees asks your policy routed to
ask, on every surface except thepathandexternal_directoryfamilies. Anything else is untouched. - For each reviewed ask it renders the structured ask facts (surface, tool names, the decision value, matched pattern, executed unit, requester provenance) into a prompt, followed by the guidance files selected for this ask. Tool results and file contents never reach the judge, and the judged value and the guidance are delimited as data, not instructions.
- The judge model must answer through a forced
report_verdicttool call (allow / deny / defer), so there is no free-text parsing to get wrong. The call is aborted aftertimeoutMs. - The verdict returns to the engine uncapped; the engine's own
bounded-delegation checkpoint enforces the
path/external_directoryline. - A circuit breaker opens after 3 consecutive model-call failures or timeouts; asks then defer instantly for a 60 second cooldown, so a down model never stalls your session.
- In-process subagents are handled correctly: registration is session-scoped, so the link always serves the session that raised the ask.
Known issues
- The classifier trusts the sessionId on the
permissions:readypayload. A stray late ready that carries a previous session's id after a newsession_startwould register the link on the old session's still-published service. This is inherent to the pi-permission-system 27.0.0 contract and the reference implementation has the same property.
Development
npm install
npx tsc --noEmit # typecheck
npx vitest run # 200 testsThe @gotgenes/pi-permission-system dev dependency is a file: reference
to a local checkout of the permission system source, so types track the
current API rather than a published copy. Adjust the path in package.json
to wherever your checkout of gotgenes/pi-packages lives.
License
MIT
