@kodel-ai/runner
v0.2.2
Published
Kodel self-hosted pull-runner daemon (block-7)
Readme
@kodel-ai/runner
Self-hosted pull-runner daemon for Kodel pipelines (block-7).
The runner is a passive executor: it declares labels (what is installed on this VM) and
opens an outbound WebSocket to the Kodel server (works through NAT). The server matches
a pipeline node to this runner and sends a runner.claim carrying a structured
AgentTask (engine + prompt + env + optional git/shell). The daemon prepares
a workspace for the node, optionally git clones and checks out the run branch, then runs
the task with the engine adapter for task.engine, streams stdout/stderr back, saves
whatever the agent produced to the repository, and reports the exit code together with what
was saved.
Formula: the server says WHAT (intent), the runner's adapter does HOW. Built-in adapters:
agentapi(claude-code / opencode via coder/agentapi) andshell(escape hatch that runstask.shell.script). There are no server-side profiles/recipes.
Install
npm i -g @kodel-ai/runnerThis installs the kodel-runner daemon.
Configure
Create a registration on the server first:
kodelctl runner create backend-1 --kind self-hosted --label claude --label agentapi --project PLT
# → prints the backing service-account PAT (token) — copy it once
# labels are assigned server-side (the daemon does not send them).Then on the VM, ~/.kodel/runner.yaml:
server: https://kodel.example.com
token: pat_xxxxxxxxxxxxxxxx # backing SA PAT (resolves the runner; never logged)
slug: backend-1 # optional; server resolves by token
workdir: ./work # node workspaces live under work/<project>/<runId>/<node>-<attempt>
# workspaceTtlHours: 24 # how long to keep them before sweeping (0 disables)
# llmAuth: proxy # proxy | personal | personal-direct (see "LLM access modes")
# concurrency: 1 # fixed at 1 in MVP (sequential)
# labels are NOT set here — they are assigned server-side at `runner create --label`.Every field can be overridden by env (KODEL_RUNNER_SERVER, KODEL_RUNNER_TOKEN,
KODEL_RUNNER_WORKDIR, KODEL_RUNNER_SLUG, KODEL_RUNNER_WORKSPACE_TTL_HOURS,
KODEL_RUNNER_LLM_AUTH) or CLI flags. Precedence (later wins): defaults < file < env < flags.
Git: the daemon owns the run branch
The daemon — not the agent — keeps the run branch. Before the engine starts it clones the
task's repository through the server's git-proxy relay and checks out the run branch; when
the engine is done it commits whatever changed (as Kodel Agent <[email protected]>, message
<TASK-KEY> <node-slug>) and pushes it. The push criterion is the state of the branch,
not who made the commit: if the agent committed on its own, the work is still pushed.
Consequences worth knowing:
- Skills must not mention git. The agent has no repository credentials (they never
leave the server) and no reason to run git commands — see
examples/skills/self-hosted/. - Nothing is lost when a push fails. The node is reported as failed with the branch and commit that stayed on the runner, the workspace is kept, and the next claim of the same run re-attempts the push before doing anything else.
- Workspaces are per run and attempt (
work/<project>/<runId>/<node>-<attempt>) and are swept by age, so a crashed node leaves evidence behind instead of being wiped by the next one.
Run
kodel-runner start --server https://kodel.example.com --token pat_xxx
# or, with the config file in place:
kodel-runner startThe daemon registers, heartbeats every 30s, and waits for claims. Reconnects with exponential backoff (1s → 30s cap). Ctrl-C to stop.
WebSocket transport
The runner uses the native ws client (NOT socket.io). This matches the Kodel server,
whose gateways are native-ws (WebSocket.Server({ noServer: true }) + an HTTP upgrade
handler on a fixed path — see tui.gateway.ts). The daemon connects to <server>/ws/runners
(http→ws, https→wss).
Auth is register-first: immediately after the socket opens, the daemon sends a
runner.register message carrying the PAT (token) + slug + labels as its first frame;
the server resolves the runner by token+slug and replies with runner.registered. The
token is sent only inside that frame (never as a query param / header) and is never logged.
Protocol messages are defined in @kodel-ai/shared ("Self-Hosted Runner Protocol"):
server → runner runner.claim / runner.cancel; runner → server runner.register /
runner.heartbeat / runner.step / runner.node.complete / runner.node.error.
Credentials
The daemon does not store any long-lived engine credential. Per-claim env (including
e.g. ANTHROPIC_BASE_URL / ANTHROPIC_AUTH_TOKEN minted by the server under the runner's
service account) arrives inside AgentTask.env and is used only for that task.
LLM access modes
Where the engines get their LLM access is a daemon setting — the server always sends the full set of credentials, and the daemon decides what the engine gets:
| Mode | Set with | What the engine uses |
| ----------------- | ---------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| proxy (default) | --llm-auth proxy | the org's ai-proxy credentials (billed to the org) |
| personal | --llm-auth personal | this machine's own subscription (claude login), routed through the Kodel gateway: calls are logged and attributed to the session owner, not to org budgets |
| personal-direct | --llm-auth personal-direct | this machine's own subscription straight to the vendor, bypassing Kodel (no LLM call logs) |
The same values go to env KODEL_RUNNER_LLM_AUTH or llmAuth: in runner.yaml.
personalrequiresclaude loginon the runner machine (as the OS user the daemon runs as). The daemon points claude-code at<server>/llm/_personaland passes the session key inANTHROPIC_CUSTOM_HEADERS(X-Kodel-Key), appended to any headers you already set there. Onlyclaude-codegoes through the gateway today;codex/opencodeinpersonalbehave likepersonal-direct.- If personal subscriptions are disabled on the Kodel instance
(
llmGateway.personalSubscriptions.enabled: false) or the server sent no gateway coordinates, apersonalsession/node fails withLLM gateway unavailable: <reason>— it never silently falls back to going direct. - Deprecated alias:
--local-llm-auth/KODEL_RUNNER_LOCAL_LLM_AUTH=1/localLlmAuth: true=personal-direct. An explicitllmAuthon the same layer wins; contradicting values on the same layer (e.g.--llm-auth personal --local-llm-auth) stop the daemon at startup. - In every mode the MCP channel (
KODEL_MCP_URL) and telemetry (OTEL_EXPORTER_OTLP_ENDPOINT) are rewritten to the daemon's--server, and the serviceKODEL_LLM_*variables never reach the engine. - The interactive terminal of a session gets the same env as the chat engine. It is
handed to tmux through a
0600file in<workdir>/.kodel/(read and deleted by the shell before it is applied, never on the command line), and tmux runs on its own socket (tmux -L kodel-runner), so a tmux server you started yourself — with your ownANTHROPIC_API_KEYin its environment — is never inherited. Anything in that tmux server's global environment that the engine should not get (e.g. left there by another daemon of the same OS user) is explicitly unset. A terminal left on the default tmux socket by a pre-KDL-477 daemon is killed when that session's terminal is opened again.
Engine adapters
The daemon picks an adapter by task.engine:
- agentapi (
engine: claude-code | opencode | codex) — drives theagentapiHTTP wrapper: spawnsagentapi server --type=<agent> -- <agent>in the checkout dir, POSTs the prompt, polls/statusuntilstable, streams the answer.agentapi(and the engine binary, e.g.claude) must be pre-installed on the VM; declare the matching labels (agentapi,claude) on the runner.--typeis required: without it agentapi guesses the engine from the binary name and garbles message parsing. Requires anagentapibuild that accepts--type(present since the flag was introduced upstream). The flag is passed unconditionally and is not probed for: falling back silently would restore the very parsing corruption it exists to prevent. If nodes start failing right after an upgrade, checkagentapi --helpfirst. - shell (
engine: shell) — escape hatch: runstask.shell.script[]sequentially, stopping at the first non-zero exit.
Reporting back to Kodel
An engine that only streams stdout leaves a step looking like exit 0 and attaches
nothing to the task. So when the server supplies KODEL_MCP_URL + KODEL_MCP_TOKEN, the
daemon wires the engine to Kodel's MCP endpoint, each in its native form:
| Engine | Wiring | Where the secret lives |
| ------------- | ---------------------------------------------------------------------------- | ------------------------ |
| claude-code | config file + --mcp-config <path> (passed last — the flag is variadic) | in the file, mode 0600 |
| opencode | OPENCODE_CONFIG_CONTENT env + PWD=<checkout> | env only |
| codex | CODEX_HOME + config.toml with bearer_token_env_var | env only |
The config file is written to <workdir>/.kodel/, outside the checkout — the checkout
is a git clone, and a config left inside it (with a token, for claude-code) could be
committed by the agent. It is removed once the node finishes.
The daemon rewrites the MCP URL to its own --server, exactly as it already does for the
ai-proxy base URLs: the server sits behind NAT and does not know the address the operator's
machine reaches it at. --llm-auth personal|personal-direct does not disable this
channel — that setting is about LLM credentials, while artifacts and step summaries are a separate concern.
Through this channel the agent can attach task files and call pipeline_report_node to
leave a summary of what it did. The token is scoped to its own run: four operations, one
task, and it dies with the run. Failure to wire the channel is not fatal — the node simply
runs as before.
Development
bun test # unit tests (config + executor)
bun run start # run the daemon locally