dsh-llama-models
v0.2.0
Published
Manage llama-server model load/unload from a dsh Settings section. Talks to the llama-server model-management endpoints directly from the browser; the configured base URL persists in the dsh settings document.
Readme
dsh-llama-models
A DeepSeek Harness (dsh) plugin that adds a
Llama Server section to the dsh web settings, so you can list, load, and unload
models on a running llama-server
without leaving the browser.
The section talks to the server's model-management endpoints directly over HTTP from the browser — no host process or RPC bridge is required.
What it does
- Lists every model the server knows about, with a live loaded / unloaded dot.
- Loads a model into VRAM with Load (→
POST /models/load). - Unloads a model to free VRAM with Unload (→
POST /models/unload). - Live status feed, like llama-server's own UI: the section subscribes to
GET /models/sseand updates the rows in real time (including load progress and failed loads), reconnecting automatically on drops. - After Load / Unload the state updates automatically: the POST answers
immediately while the real load/unload runs in the background, so the button
shows … until the status feed reports the model actually reached
loaded/unloaded— then the list is re-read once. If the server has no/models/sseendpoint, the wait falls back to pollingGET /models(60 s deadline). - Connect points the section at any llama-server by base URL; the URL is a
durable preference, written to the dsh settings document
(
~/.dsh/settings.yaml, underllama-models: { baseUrl: … }) and restored on every open. A remote browser (where the settings RPCs are loopback-only) falls back tolocalStorage. - ↻ re-reads the model list.
Requirements
Enable CORS on llama-server
The dsh web UI runs on one origin and the server on another, so the browser enforces
CORS. llama-server only sends a usable Access-Control-Allow-Origin header when you
pass --cors-origins. Without it the request is silently blocked and the section
shows a "Cannot reach llama-server" error.
# allow the dsh web UI origin (recommended for a single known client):
llama-server --cors-origins http://127.0.0.1:3080
# or allow any origin (matches llama.cpp's default; insecure without an API key):
llama-server --cors-origins *If you set an API key, scope the allowed origins explicitly instead of *:
llama-server --api-key sk-... --cors-origins http://127.0.0.1:3080Endpoint compatibility
Uses the llama.cpp model-management API:
| Method & path | Body | Purpose |
| --------------------- | ------------- | ------------------------------------ |
| GET /models | — | list models and loaded status |
| POST /models/load | { "model": "id" } | load a model into VRAM |
| POST /models/unload | { "model": "id" } | unload a model to free VRAM |
| GET /models/sse | — | live status feed (server-sent events) |
GET /models returns { "object": "list", "data": [ { "id", "status": { "value":
"loaded" | "unloaded" }, "source", ... }, ... ] }. The loaded state is read from
status.value (a top-level loaded boolean is accepted too, for older builds).
The load/unload endpoints answer immediately (HTTP 200 {"success": true}); the
real work runs in the background and is announced on GET /models/sse as
status_change records (data.status = loading / loaded / unloaded /
failed, with exit_code on unload and optional progress while loading).
Build
pnpm install # from the workspace root
pnpm --filter dsh-llama-models bundle # -> lib/client.jstsdown emits a CommonJS bundle to lib/client.cjs;
scripts/build-client.mjs wraps it in the
window.__ModuleLoader__.load({ id, factory }) shape the dsh browser expects and
publishes it as lib/client.js.
Install to your local dsh
The compiled bundle installs as a dependency of the web profile
($DSH_HOME/profiles/web, usually ~/.dsh/profiles/web). Discovery is
two-staged: at boot, each entry in the profile's dsh.profile.bundles layer
stack is loaded by the profile loader, and the web app's client-module
service scans those loader entries for packages that declare dsh.client
(platform: "web") and export ./client. A plain node_modules dependency
is never scanned — which is why this package also declares dsh.bundle with
cordis.patch.yml, whose insert row puts the plugin
into the profile roster. dsh plugin add then reconciles the profile
automatically: because the package declares dsh.bundle, it is appended to
dsh.profile.bundles in the profile's package.json with no manual editing.
Two packaging consequences of the roster row:
- The loader imports every roster entry as a Node module at boot and
expects a cordis plugin (a function or an object with
apply). This is a pure browser plugin, soindex.jsis a no-opapplyexposed throughmain/exports["."]— without it the profile fails to boot withERR_PACKAGE_PATH_NOT_EXPORTED. - The browser half is picked up separately:
dsh.client(platform: "web") plusexports["./client"]pointing at the builtlib/client.js.
1. Pack the compiled result
cd packages/llama-models
pnpm pack # -> dsh-llama-models-0.1.0.tgzThe tarball honors the files field: it contains only lib/client.js, the
type declarations, cordis.patch.yml, and package.json — no source.
2. Install into the web profile
dsh plugin --profile web add --store-dir ~/.dsh/.pnpm-store \
./dsh-llama-models-0.1.0.tgzdsh plugin runs pnpm in the profile directory and records the tarball as a
file: dependency in the profile's package.json. (The ./ prefix matters:
dsh plugin anchors relative path specs against your invoking directory, while
pnpm itself runs with the profile directory as its cwd.)
On this machine the --store-dir flag is required for every dsh plugin
subcommand (add and remove): the profile's node_modules was linked from
~/.dsh/.pnpm-store, while pnpm 11's default store is
~/.local/share/pnpm/store — without it, pnpm fails with
ERR_PNPM_UNEXPECTED_STORE.
A successful install should be silent about bundle warnings, and the profile's
package.json should now list dsh-llama-models under
dsh.profile.bundles. If you see the declares no dsh.bundle warning, the
tarball predates the dsh.bundle declaration — repack and re-add. You can
dry-run the composed tree with dsh --profile web --dump-config and grep for
llama-models without booting the app.
3. Restart dsh web
Client bundles are loaded at web boot, so restart the running dsh web
instance to pick the plugin up, then open ⚙ Settings → Llama Server.
Updating
After changing the source and rebuilding — still from this package directory:
npm run bundle # recompile -> lib/client.js
pnpm pack # re-create the tarball
dsh plugin --profile web remove --store-dir ~/.dsh/.pnpm-store \
dsh-llama-models
dsh plugin --profile web add --store-dir ~/.dsh/.pnpm-store \
./dsh-llama-models-0.1.0.tgz # re-import the new snapshotThe remove + add pair is deliberate: the tarball is a snapshot, and a
fresh add guarantees the profile re-imports the new contents. Then restart
dsh web again.
Dev alternative — live directory link: install the package directory instead of a tarball, so every rebuild is picked up without re-installing:
dsh plugin --profile web remove --store-dir ~/.dsh/.pnpm-store dsh-llama-models
dsh plugin --profile web add --store-dir ~/.dsh/.pnpm-store "file://$(pwd)"then just re-run npm run bundle after each change. To go back to the
tarball form, remove the dependency and repeat the install steps above.
Usage
- Start llama-server with
--cors-origins(see above). - Install the built plugin into the web profile (see Install to your local dsh).
- Open ⚙ Settings → Llama Server.
- Enter the server base URL (e.g.
http://127.0.0.1:8008) and press Connect. - Use Load / Unload to move models in and out of VRAM.
Project layout
src/client.tsx # React Settings section + llama-server HTTP client
tsdown.config.ts # tsdown build (CJS, browser platform)
scripts/build-client.mjs # wraps the bundle for the dsh module loader
cordis.patch.yml # bundle patch: inserts the plugin row into the profile roster
index.js # node entry: registers the `llama-models` settings namespaceThe node half is a real cordis plugin: it registers the llama-models settings
namespace (a baseUrl string field) with the host's settings provider, which is
what makes the base URL durable. The browser half reads and writes that namespace
through ctx.settingsScope — Connect persists the URL, and the input re-seeds
from it on every open. It calls llama-server directly for the model list and
load/unload, so it needs only the slots and settingsScope cordis services to
register in the settings surface. If the settings service is absent (e.g. a minimal
profile), the section degrades to localStorage for the URL.
