@localaicat/pi
v0.1.6
Published
pi extension that warms/cools the Local AI Cat resident model over the Local API
Maintainers
Readme
@localaicat/pi
A pi extension that keeps the Local AI Cat resident model warm for the duration of a coding session and publishes local generation stats to pi's status line. It also exposes resource-control slash commands so you can load, unload, or stop the local server without a custom pi binary.
The Local API loads a model lazily on the first chat request, so the first turn
of a session is slow (weights load + KV warm-up). This extension proactively
loads the model when pi selects one from the localaicat provider, and
releases its hold on session shutdown. The model is shared: other pi
sessions and clients may be using it, so the app drops it only when the last
holder releases (see Shared model holds).
Install
# from npm (recommended)
pi install npm:@localaicat/pi
# or install a local checkout persistently
pi install ./tools/pi
# or one-off for a single run
pi -e ./tools/pi/index.tsQuick start
pi install npm:@localaicat/pi
export LOCALAICAT_BASE_URL=http://127.0.0.1:11434
export LOCALAICAT_API_KEY=local # if the Local API requires a token
export LOCALAICAT_MODEL=mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit
pi --provider localaicatThe extension self-configures the localaicat provider on first load by
writing ~/.pi/agent/models.json from the LOCALAICAT_* environment above
(see "Provider self-configuration"), so the install + four lines is all you need.
Provider self-configuration
On load the extension ensures pi knows about the localaicat provider. If
~/.pi/agent/models.json has no localaicat entry, it writes one:
{
"providers": {
"localaicat": {
"baseUrl": "http://127.0.0.1:11434/v1",
"api": "openai-completions",
"apiKey": "local",
"models": [{ "id": "mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit" }]
}
}
}baseUrl comes from LOCALAICAT_BASE_URL (with /v1 appended), apiKey from
LOCALAICAT_API_KEY (default local), and the seeded model from
LOCALAICAT_MODEL (default Qwen3-Coder-30B-A3B-Instruct-4bit). An existing
localaicat entry is never overwritten — hand-edit it freely. If the write
fails for any reason, paste the block above manually.
Environment
| Variable | Default | Purpose |
|---|---|---|
| LOCALAICAT_BASE_URL | http://127.0.0.1:11434 | Base of the Local API |
| LOCALAICAT_API_KEY | – | Local API token when the app requires one |
| LOCALAICAT_PROVIDER | localaicat | Comma-separated pi provider names to match on model_select |
| LOCALAICAT_MODEL | – | Model to warm on session_start if none selected yet |
| LOCALAICAT_LOG | – | Append one line per load/unload action (used by tests) |
| LOCALAICAT_REQUEST_TIMEOUT_MS | scaled | Request wait, exactly (0 = no limit). See Request wait |
| LOCALAICAT_PREFILL_TOKENS_PER_SEC | 25 | Prefill rate the scaled wait assumes |
| LOCALAICAT_UNLOAD_ON_EXIT | – | always / never override for releasing on exit (see below) |
Request wait
A long prompt is read (prefilled) before the model sends its first byte: 14,098
tokens took 234 s on Qwen3.8-27B-4bit before the engine fix. pi's own request
wait is its httpIdleTimeoutMs setting (5 min by default), after which pi used
to retry three times, each retry reading the whole prompt again.
For localaicat models the extension now waits
max(pi's setting, clamp(5 min + estimated prompt tokens / 25 tok/s, 10 min, 2 h)).
LOCALAICAT_REQUEST_TIMEOUT_MS replaces that with an exact value. If the wait
still runs out, pi shows "Local AI Cat sent no reply in N min while reading a
prompt of about Nk tokens…" and does not resend it on its own; send the
message again to wait once more.
Shared model holds
Each pi process has a client id (a UUID, stable across /reload). Load and
unload send it as the x-localai-client-id header and as {"client_id": "…"}
in the JSON body; chat requests carry the header too. Load takes a hold; unload
releases this client's hold, and the app unloads only when no holder is left.
An app that tracks holds answers load/unload with client_id (echoed) and/or
holders (clients still holding). If the app's load answer has neither, it
predates holds, and an unload would pull the model from under every other
client, so on shutdown the extension leaves the model resident (the app's idle
auto-unload still frees it). LOCALAICAT_UNLOAD_ON_EXIT=always restores the old
unload; never never releases. An explicit /localaicat unload is always sent.
Local Coding in the app
Loading a model turns the app's Local Coding on for it. When the app
advertises it ("prepare_accepts_model": true on GET /v1/local/coding/status),
the extension loads through POST /v1/local/coding/prepare with
{"model": "<id>", "client_id": "<uuid>"}. That loads on the engine that serves
the model, records this client's hold exactly like /models/{id}/load, and makes
the model the app's selected Local Coding model, so the menu bar reads "Purring"
for the model pi is running. Shutdown releases the hold as before; the menu bar's
own on/off takes a separate hold, so turning Local Coding off in the menu does not
unload a model pi still holds, and pi quitting does not unload one the menu holds.
Older apps (no marker, or a prepare that refuses the model — e.g. a model outside
the app's Local Coding list) get the plain POST /v1/local/models/{id}/load.
The status probe exists because an older app's prepare ignores the body and would
load its own selected model instead.
Commands
All under a single /localaicat command:
/localaicat(or/localaicat status) — show loaded state + the local RAM budget / headroom (/v1/local/resources)./localaicat load [model-id]— load (and turn Local Coding on for) a specific model, else the model this session warmed, else pi's selectedlocalaicatmodel, elseLOCALAICAT_MODEL./localaicat unload [model-id]— release this session's hold on the same target./localaicat stop— stop the Local API server.
Status line
The extension updates pi's footer with the last assistant turn when pi exposes usage data:
localaicat: loaded Qwen3-Coder-30B-A3B-Instruct-4bit · 28.4 tok/s · 1240/310/1550 tok · 12.8% ctxWhile a request is in flight it reads busy <model> instead (the resource
snapshot can miss a model that is serving). Otherwise it reads loaded/unloaded
state from /v1/local/resources, OpenAI-style usage
fields (prompt_tokens, completion_tokens, total_tokens,
tokens_per_second), and pi's built-in context estimate via
ctx.getContextUsage(). If a value is missing, it is omitted rather than
blocking the session.
How it works
| pi event | action |
|---|---|
| model_select (provider = localaicat) | POST /v1/local/coding/prepare (turns Local Coding on), else POST /v1/local/models/{id}/load |
| agent_start / message_update | footer shows busy <model> (no polling) |
| message_end | repair wrapped tool-call arguments; footer resource + stats |
| agent_end | footer back to resource state |
| session_start (with LOCALAICAT_MODEL) | warm fallback load |
| session_shutdown | POST /v1/local/models/{id}/unload (release this client's hold) |
| /localaicat stop | POST /v1/local/server/stop |
If the app isn't running, actions degrade quietly — the next chat request still loads the model lazily.
Tool calls whose arguments arrive one level too deep
({"arguments": "{\"command\": …}"}) are unwrapped before pi validates them.
Tests
npm test # unit + extension tests (node --test)
npm run check # typecheck
npm run test:real-pi # real pi binary vs a fake Local API (PI_BIN=… to choose)