@nanobpm/nano-coder
v0.42.0
Published
A 6MB coding agent. Run a fleet on your laptop.
Readme
nano-coder
A 6MB coding agent. Run a fleet on your laptop.
nano-coder is a coding agent for the terminal, written in Rust. It uses about 6MB of resident memory, where Node-based agent CLIs take 150–660MB, so you can run dense fleets of agent workers on one machine. It runs interactively, or headless over ACP as a worker for nano-workforce via c8ctl-nano.
Install
npm install -g @nanobpm/nano-coder # prebuilt binaries for macOS and Linux (x64, arm64)
cargo install nano-coder # or build from sourceFeatures
- Interactive CLI: REPL-based interface for conversing with the agent
- ACP Protocol: JSON-RPC 2.0 over stdio for headless orchestration (c8ctl-nano compatible)
- Providers: OpenAI-compatible and Anthropic endpoints (remote or local), selected per model as
provider/model, with retry/backoff - Tool Calling: Agent can invoke registered tools during conversation, including a real
bashtool with timeouts and bounded output - Sessions: Append-only JSONL session logs with resume and input-ID deduplication
- Lifecycle Hooks: 6 internal hook events for observing agent behavior, plus
Claude Code–compatible external user hooks loaded from
.claude/settings.json - Configuration: TOML-based config file at
~/.config/nano-coder/config.toml - Commands:
/help,/compact,/context,/verbosity,/settings,/tools,/skills,/queue,/restart,/exit - Streaming output: answers stream in, thinking shows collapsed (Ctrl-O expands it), tool calls show inline
- Status line pinned to the bottom of the terminal, plus manual and automatic context compaction
- Task plans:
plan_*tools keep a plan with notes outside the conversation, so long tasks survive compaction, resume and a change of worker - Project instructions:
AGENTS.md(orCLAUDE.md,.github/copilot-instructions.md) from the repository is added to the system prompt - Safety: built-in guards block destructive commands (
rm -rf /,DROP DATABASE, force-pushingmain, ...), user allow/deny rules, and an optional OS sandbox (Seatbelt on macOS, Landlock on Linux) - Skills:
SKILL.mdfolders from the repository,~/.agents/skills, and an spmai.lock, loaded on demand withload_skill
Two Execution Modes
Interactive CLI Mode (default)
nano-coderStarts the interactive REPL where you can chat with the agent and use slash commands.
While a turn is running you can type a message and press Enter to steer the agent: the message is added to the conversation before the agent's next step (after the tool calls in flight finish), so it can change course without the turn being cancelled. A steer that arrives as the turn finishes is queued instead, and one that arrives as the turn is cancelled is dropped (with a note showing its text).
Press Ctrl-Enter (or Cmd-Enter) instead to queue the message: it waits in a queue
and runs as a later prompt — one per turn, in the order you sent them. Ctrl-Enter needs a
terminal that reports modified keys (modifyOtherKeys or the kitty keyboard protocol, e.g.
xterm, kitty, WezTerm, Ghostty, iTerm2); elsewhere use /queue add message. The status
line shows how many are waiting. Edit the queue at any time
(even mid-turn) with /queue: /queue lists the waiting messages, /queue remove N
drops one (or several, /queue remove N M), /queue edit N new text rewrites one, and
/queue clear empties the queue. A queued message is never injected while the agent is
asking for input (the question tool's picker owns the terminal until you answer).
Slash commands work mid-turn too, where it is safe:
- Right away: read-only ones (
/help,/tools,/skills,/providers,/session,/plan,/context,/trajectory(--json/--markdown), and/modeor/verbositywithout an argument) show the current state./planshows the plan as the agent updates it. - From the agent's next step:
/mode NAMEapplies when the agent takes its next step. - Right away (for the rest of the turn):
/verbosity LEVELchanges the output level immediately, so it affects the remaining events streamed by the current turn. - After the turn: commands that change the conversation or take over the keyboard
(
/compact,/restart,/settings,/model,/exit) wait for the turn to finish.
Esc Esc (twice within a second) or Ctrl-C cancels the running turn, killing any running bash command; a second Ctrl-C at the prompt exits. At the idle prompt, Esc Esc instead clears the input, so a half-typed or pasted prompt can be discarded without deleting it character by character. With piped (non-terminal) stdin, lines read during a turn are queued as later prompts (never steer).
ACP Headless Mode (--acp flag)
nano-coder --acpSpeaks the Agent Communication Protocol (ACP) over stdio using newline-delimited JSON-RPC 2.0 messages. Compatible with c8ctl-nano's spawnCaptureAcp executor.
Protocol methods supported:
initialize→ returns protocol version and capabilities (loadSessionwhen persistence is on)session/new→ starts a fresh conversation (and session log), returns sessionId.params.cwd(absolute, existing directory) becomes the working directory for toolssession/load→{ "sessionId": ... }resumes a persisted session, first replaying the conversation assession/updatenotifications (as the ACP spec requires)session/prompt→ processes a prompt, supports tool calls. One session is active per process; asessionIdother than the active one is rejected (usesession/loadto switch). An optionalmessageId(or_meta.inputId) makes it idempotent: redelivering an ID that already completed returns the recorded response without calling the model againsession/cancel→ cancels the current turn (normally sent as a notification, which gets no reply). The running model call is abandoned, a running bash command is killed, remaining tool calls are recorded as cancelled, and the turn'ssession/promptresolves withstopReason: "cancelled"
Steering. A session/prompt for the active session that arrives while a turn is
running is a steer, not a queued prompt. It is added as a user message before the next
model call (echoed as a user_message_chunk) and answered when the turn ends with
{ "stopReason": ..., "_meta": { "steered": true, "inputOf": <main prompt id> } }.
A steer that arrives too late to join the turn runs as an ordinary prompt, or is
answered with stopReason: "cancelled" if the turn was cancelled. Slash-command prompts
and other requests that arrive mid-turn are handled after the turn ends. Everything
handled after the turn, late steers included, runs in the order the client sent it.
stopReason is end_turn, cancelled, or max_turn_requests.
Streaming updates. During a turn the harness sends session/update notifications:
agent_message_chunk (with a messageId), tool_call (toolCallId, title = tool name,
kind, rawInput) and tool_call_update (completed/failed, rawOutput). These are the
events c8ctl-nano's transcript producer records into the engine's AgentInstance history.
Project instructions. session/new and session/load results include
_meta.projectInstructions: the absolute paths of the instruction files that were added to
the system prompt for the session's cwd (see Project Instructions).
_meta.skills lists the skills found for it, and _meta.skillWarnings (when present) explains
any that could not be loaded (see Skills).
Plans. Each plan change sends a plan update: ACP entries (content, priority,
status) plus _meta.plan, the full plan with ids, notes and dependencies. To continue a job
on another worker, pass the last _meta.plan as session/new params._meta.plan (bare ACP
entries are accepted too). The new session starts with that plan, the session/new result
echoes it in _meta.plan, and the model is told about it with the first prompt (unless
_meta.planInPrompt: true says the client already put the plan in the prompt). initialize
advertises this as agentCapabilities._meta.planSeed. See Task Plans.
Outcomes. When the model calls report_outcome, the session/prompt result carries
_meta.outcome: {"status": "completed" | "blocked" | "needs_input", "summary": "..."}. A client can use it
instead of guessing from the stop reason: blocked means the model needs help (an
escalation). Redelivering the input returns the same outcome. See Outcomes.
Slash commands work via ACP too:
/compact [--smart|--standard] [focus]- summarizes the conversation; the result hascompacted,before,after,tokensBefore,tokensAfter,summarized,mode,fallbackandtruncated/settings- returns current settings as JSON/tools- lists registered tools/plan- returnsplan(JSON) andtext(the rendered plan)/providers- lists providers/model provider/model- switches model
Exit status. Normal shutdown — stdin closing after at least one valid request was
processed — exits 0. If stdin reaches EOF without a single parseable JSON-RPC request
(e.g. the client isn't speaking ACP), the harness prints
no valid ACP requests received on stdin … to stderr and exits 2, so a misconfigured
caller can't mistake a no-op run for success. A non-JSON first line is called out
explicitly (input doesn't look like ACP JSON-RPC …).
Architecture
src/
├── main.rs # Single binary entry point (interactive + ACP modes)
├── agent.rs # Agent core: conversation management, tool execution loop
├── acp.rs # ACP JSON-RPC protocol handler
├── hooks.rs # Internal (observe-only) lifecycle hook registry and events
├── claude_hooks.rs # External Claude Code–compatible user hooks engine
├── tools.rs # Tool registration and dispatch system
├── llm.rs # Provider-neutral messages and the async LLMClient trait
├── providers/ # Provider registry + presets, HTTP transport with retries
│ ├── openai.rs # OpenAI Chat Completions (and compatible servers)
│ ├── anthropic.rs # Anthropic Messages API
│ ├── github_copilot.rs # UNOFFICIAL Copilot-subscription provider
│ ├── retry.rs # Retry classification and backoff
│ └── mock.rs # Offline scripted client
├── bash.rs # bash tool: timeout, file capture, bounded output
├── files.rs # read_file / write_file / edit_file tools
├── shell.rs # Bash parser used by the permission checks
├── permissions.rs # Allow/deny rules and built-in guards against destructive commands
├── sandbox.rs # Seatbelt (macOS) / Landlock (Linux) sandbox for shell commands
├── output.rs # Head/tail output bounding, spilling long output to disk
├── session.rs # Versioned append-only JSONL session log
├── session_index.rs # Session summaries (.index.jsonl) for the --resume picker
├── resume.rs # --resume picker, --resume last, --list-sessions
├── context.rs # Token accounting, context-window heuristics, overflow detection
├── status.rs # Bottom-of-terminal status line
├── ui.rs # Verbosity levels and the streaming output renderer
├── frame.rs # App-owned frame renderer (renderer = "frame"): full redraw on resize
├── lineedit.rs # Key-by-key prompt input (Ctrl-O, mid-turn input on the status line)
├── input_history.rs # Up/Down recall of submitted lines (session-scoped)
├── instructions.rs # AGENTS.md / CLAUDE.md discovery for the system prompt
├── plan.rs # Task plan and the plan_add / plan_update / plan_show tools
├── queue.rs # The interactive message queue and the /queue editor
├── goal.rs # report_outcome tool (completed / blocked / needs_input)
├── memory.rs # Cross-session memory: memory_save / memory_search / memory_forget
├── mode.rs # Agent mode (normal / plan / auto) and the plan-mode tool gate
├── question.rs # question tool and the mid-turn question / turn-cap rendezvous
├── commands.rs # Slash-command table for /help and the as-you-type menu
├── skills.rs # SKILL.md discovery, ai.lock sources and the load_skill tool
├── reminders.rs # <system-reminder> notes appended to tool results
├── settings.rs # /settings menu and config-file writer
└── config.rs # Configuration file loading and managementMode selection: --acp flag enables ACP headless mode; default is interactive CLI.
Lifecycle Hooks
nano-coder has two hook systems.
Internal hooks (observe-only)
The harness exposes 6 internal lifecycle hook events for in-process Rust
callbacks. These are observe-only: a callback can inspect the event payload
(for logging, metrics, debugging) but cannot block, modify, or redirect the
agent. They are registered programmatically against agent.hooks().
| Hook | When it fires |
|------|---------------|
| before_context_load | Before processing user input |
| after_context_load | After adding user message to conversation |
| before_llm_send | Before sending messages to LLM |
| after_llm_response | After receiving LLM response |
| before_tool_call | Before executing a tool |
| after_tool_call | After tool execution completes |
External user hooks (Claude Code–compatible)
nano-coder also runs external user hooks that follow the Claude Code hooks protocol, so existing Claude Code hook configurations work unchanged. These are loaded, in order, from:
[hooks]in nano-coder's own user config (config.toml), in Claude's structure~/.claude/settings.json(your Claude Code user settings; disable withclaude_user_hooks = false)<project>/.claude/settings.json(project settings, committed)<project>/.claude/settings.local.json(project-local, git-ignored)
Each hook is a command that nano-coder runs as a subprocess, passing the
event payload as JSON on stdin and interpreting its exit code and (optional)
JSON stdout per Claude's protocol. Hooks can block a tool call or prompt,
modify a tool's input, or inject additional context. nano-coder
translates its tool names (bash→Bash, read_file→Read, write_file→
Write, edit_file→Edit, grep→Grep, glob→Glob) and the file_path/
path argument so Claude-authored matchers and scripts match correctly.
Supported events: SessionStart, UserPromptSubmit, PreToolUse,
PostToolUse, and Stop. A PreToolUse hook that fails (crashes, times out,
or exits non-zero) fails closed and blocks the tool call; failures in other
events are non-blocking.
Run /hooks at the prompt to list the loaded hooks, which files they came
from, and any that were skipped (e.g. unsupported events reserved for later
phases). Disable them with --no-hooks, or granularly via the
disable_hooks, disable_project_hooks, and claude_user_hooks settings.
Built-in Tools
get_time- Get current date and timeecho- Echo back input textbash- Runbash -c <command>(stdin closed, own process group). Arguments:command, optionaltimeout_seconds(defaultbash_timeout_secs, 600) andmax_output_length(default 40,000, max 1,000,000 characters each for stdout and stderr). Returns stdout, thenStderr:andExit code: Nwhen relevant, or(no output). Long output keeps its head and tail with...N bytes truncated; complete output in <path>...; the full capture stays in that file.read_file- Numbered lines of a text file;path, optionaloffset(1-based) andlimit(default 2000 lines). Refuses binary files.write_file- Create or overwrite a file (path,content), creating parent directories.edit_file- Replace text (path,old_string,new_string, optionalreplace_all). Fails unlessold_stringmatches exactly once (orreplace_allis set). With no exact match, it retries ignoringread_fileline-number prefixes, then trailing whitespace, then indentation (re-indentingnew_stringto fit), and uses such a match only if it is unique; otherwise the error shows the closest lines. Returns the edited lines with 3 lines of context.edit_file, andwrite_fileover an existing file, require the file to have been read withread_file(or written by these tools) in this process, and to be unchanged on disk since. Otherwise they ask the model to read it again, so an edit can't land on stale text.plan_add,plan_update,plan_show- The agent's task plan (see Task Plans).report_outcome- Report the taskcompleted,blockedorneeds_input, with asummary; ends the turn (see Outcomes).question- Ask the user a structured question and wait for the answer, without ending the turn. Interactive sessions only; in auto mode an unanswered question is resolved after 15s with the away message. Headless (ACP) sessions get an error instead.load_skill- Return a skill's instructions and list its other files;name. Offered only when skills were found (see Skills).memory_save,memory_search,memory_forget- Cross-session memory: facts the model saves in one session and finds in later ones (see Memory). Offered whenmemoryis on; read-only runs (headless/ACP) offermemory_searchonly.
Any other tool's result longer than 40,000 characters is cut the same way as bash output,
with the whole result saved under the temp directory (nano-coder-<pid>/tool-<id>-<name>.txt)
and its path in the marker. read_file pages instead.
Relative paths resolve against the working directory (ACP session/new cwd). Writes are
atomic (temp file + rename). There is no permission prompt; every call is checked against
the permission rules and sandbox instead.
Permissions and Sandbox
nano-coder never stops to ask for approval (it runs headless in agent fleets). Every tool call is checked before it runs instead, and a blocked call returns an error telling the model to stop and ask the user rather than work around the block.
Order of checks: deny rules, then allow rules, then the built-in guards. Deny always
wins. An allow rule approves a shell command only when every command in it matches, so
Bash(git *) does not approve git status && rm -rf /.
Rules name a tool and an optional pattern:
| Rule | Matches |
|---|---|
| Bash(rm -rf *) | a shell command; * matches anything, including spaces and / |
| Bash(git push:*) | git push alone or with any arguments |
| Read(~/.ssh/**) | read_file paths; ** crosses directories, * does not |
| Edit(**/.env) / Write(...) | write_file and edit_file paths (relative to the working directory, or absolute) |
| write_file, bash, any tool name | every call to that tool |
Shell commands are parsed, not pattern-matched as text. The command line is split on ;,
&&, ||, |, &, newlines and parentheses; quotes are removed; $(...), backticks and
<(...) are parsed as further commands. The guards and rules then see through assignments
(FOO=1 cmd), wrappers (sudo, env, timeout, nice, xargs, nohup, command, ...),
bash -c '...', eval, ssh host cmd, find -exec, and here-documents fed to a shell. A command
that cannot be parsed, or a script computed at run time (bash -c "$CMD",
eval "$(curl ...)"), is blocked.
Built-in guards (builtin_rules = true) block:
- recursive
rm(andmv,find -delete,chmod -R/chown -R) on/, your home directory or its top-level folders, the working directory or its parents, top-level and system directories, and.git.rm -rf *counts as the working directory. An unset variable counts as empty, sorm -rf "$DIR/"*is blocked unless written"${DIR:?}/"*. Paths follow acdand variable assignments earlier in the same command (cd .. && rm -rf projectis blocked) mkfs,fdisk,wipefs, destructivediskutil,dd of=/dev/...and redirects to devices, fork bombs,shutdown/reboot- destructive SQL (
DROP DATABASE|SCHEMA|TABLE,TRUNCATE,DELETE FROMwithoutWHERE,ALTER TABLE ... DROP,FLUSHALL,dropDatabase()) in a command that uses a database client (psql,mysql,sqlite3,mongosh,redis-cli, also viadocker exec) or inline interpreter code (python -c), plusdropdb,rails db:drop,prisma migrate reset,manage.py flush terraform destroy,pulumi destroy,kubectl delete namespace|--all,aws s3 rbgit push --force(or+refspec,--all, wildcard refspecs) to, or deleting, a protected branch, andgit push --mirror
Add an allow rule for anything legitimate they block, e.g.
allow = ["Bash(sqlite3 test.db *)"], or set builtin_rules = false.
These checks catch mistakes, not adversaries. A model can write a script and run it, and nothing inspects that. The boundary is the OS sandbox, plus credentials: don't give the agent production database URLs or broadly scoped tokens.
Sandbox (--sandbox workspace, off by default) runs each shell command under Seatbelt
(sandbox-exec) on macOS or Landlock on Linux (6.2+). Commands can read everywhere, but
write only to:
workspace: the working directory, its git directories (including a worktree's shared one), temp directories, package-manager caches (~/.cargo/registry,~/.npm,~/.cache,~/Library/Caches,~/.gradle,~/go/pkg/mod, ...) andwritablepathsread-only: temp directories andwritablepaths
write_file and edit_file are held to the same directories. network = false blocks
outbound connections (macOS: except to localhost; Linux: all TCP, which needs Linux 6.7+).
If the sandbox is enabled but cannot be applied, commands fail instead of running
unsandboxed. When a sandboxed command fails with a permission error, the result tells the
model where it may write.
Commands
A line starting with / is never sent to the model. Enter on an unknown command (/exin)
keeps it in the input with a hint (Unknown command /exin. Did you mean /exit?) so you
can fix it. To send a message that starts with /, such as a path, type //:
//usr/lib is 4 GB, why? sends /usr/lib is 4 GB, why?.
Typing / at the prompt lists the commands under it, and each further character narrows the
list. Tab completes the command, or the part all matches share. Esc hides the list. The
list is built from the same table as /help (src/commands.rs). Commands with a known
argument set (/model, /mode, /verbosity, /thinking) get the same treatment for their first
argument: a type-ahead list narrows as you type and Tab completes it.
/help- Show available commands/compact [--smart|--standard] [focus]- Summarize older messages with the current model, keeping the latest message.--smart/--standardoverridecompaction_modefor this compaction (see Smart compaction). Optional text tells the summary what to focus on. Esc Esc or Ctrl-C cancels/verbosity [quiet|normal|verbose|debug]- Show or set how much is printed (see below)/context- Show context usage, window, session token totals, AI Credits (GitHub Copilot), auto-compaction state and the loaded instruction files/settings- Interactive settings menu:- Model: starts with the last four models you used, then pick a provider and a model from its live model list (or type an ID)
- Add or edit a provider: name, API kind (OpenAI-compatible, Anthropic, Copilot), base URL, key source (env var, shell command, or a literal key; the file is then written with mode 0600) and default model
- temperature, max tokens, system prompt
- Turn cap (
max_iterations): LLM calls per input (0 = unbounded); a positive cap makes normal mode ask before stopping - Context: auto-compaction on/off, threshold, compaction mode, context-window override
- Verbosity
- Save to config file: writes only the keys you changed into the config file
(
--configor~/.config/nano-coder/config.toml), keeping comments and other settings. Leaving with unsaved changes asks whether to save
/tools- List registered tools/skills- List the skills the agent can load, where each lives, and any loading warnings/plan- Show the agent's task plan with all notes/memory [forget ID]- List cross-session memories with their ids, or delete one by id (see Memory)/hooks- List the loaded external user hooks (from.claude/settings.jsonand friends), the files they came from, and any skipped entries (see Lifecycle Hooks)/queue [list|add text|remove N...|edit N text|clear]- Show or edit the queued messages. Works while a turn runs, so a queued message can be removed or rewritten before it is sent./model [provider/model]- Show the current model and pick a new one. The list starts with the last four models you used (the current one marked; the previous one highlighted, so/modelthen Enter switches back), then the providers: pick a provider to scroll its model list (Esc steps back). With an argument, switches directly (conversation is kept). Typing/modelshows a type-ahead of the current model, recently used models, and each configured provider's default model; Tab completes (a bare provider name completes to its default model). Recently used models are kept in~/.local/share/nano-coder/recent-models.json/mode [normal|plan|auto]- Show or set the agent mode (Shift+Tab cycles it, at the prompt or mid-turn):- normal - full tools; reaching a positive turn cap asks whether to keep going
- plan - read-only: mutating tools (
bash,write_file,edit_file) are gated, only analysis and output - auto - no turn cap; a
questionleft unanswered for 15s is answered with "the user is away from the keyboard, make the best decision you can"
/thinking [level|default|off|reset]- Show the thinking level in use and the levels the current model takes, or set one for this session (resetgoes back to the configured level). See Thinking/providers- List providers, endpoints and whether their API key is available/session- Show the session ID and log path/trajectory- Show this session's trajectory turn by turn: user input, thinking, answers, tool calls and results, tokens and timings. Each message row is labelled with its#Nsession-log ID, the same IDhistory_readand smart-compaction summaries use (compaction and crash-recovered input rows have no#N, as they aren't cited that way). When it doesn't fit on the screen it opens in your pager ($PAGER, defaultless) — but only at an idle prompt with the frame renderer: invoked mid-turn (while a turn runs) or under the legacy renderer it prints inline instead./trajectory --jsonor/trajectory --markdownprints an export instead;nano-coder --trajectory SESSION_ID [--json|--markdown]does the same for any saved session/resume [ID|last]- Switch to a saved session without restarting: the same picker as--resume(leaving out the session in use), or the session with that ID, orlast(the most recent other session in this directory). Run it at the prompt, not during a turn/restart- Start a fresh session (clean context) without exiting/exit- Exit the agent (prints the session's--resumecommand first, when session persistence is enabled)
Building and Running
cargo build --release
./target/release/nano-coderOr run directly:
cargo run
cargo run -- --model anthropic/claude-sonnet-4-5
cargo run -- --model ollama/qwen2.5:1.5b
cargo run -- --resume sess-20260923T012518-7e7923f8
cargo run -- --resume # pick a session
cargo run -- --resume last # the most recent session in this directoryResuming. --resume without an ID opens a picker of this directory's saved sessions, most
recently used first. Each row shows when the session was last used, its project (the last part of
its directory), how many prompts it has, and its last prompt. When the last prompt says little
("do it"), the row also shows the more telling prompt before it. Type to filter, Enter to resume,
Esc to cancel. The last entry shows sessions from every directory. At the prompt, /resume does the same without restarting. Sessions with no prompts are
left out. Sessions from before this feature don't record their directory: they are shown in every
directory, with ? as the project. Without a terminal, --resume prints the list and exits.
--list-sessions [--all] [--json] prints the list for scripts (--all: every directory).
The picker reads .index.jsonl in the session directory: one summary per session, updated at the
end of each turn. It is only a cache. A session that is missing from it, or whose log changed
since it was indexed, is summarized from its log again, and deleting the file rebuilds it.
Titles. With session_titles = true, once a session has a prompt that says something (not
just "hi"), nano-coder asks the model for a title of at most six words in the background. It uses
title_model (a cheap one is enough) or the session's model. The request is a few hundred tokens,
tried at most once per session per run (so a failed or deleted title is retried after a restart,
not on the next turn). The picker then shows title · last prompt. Titles live only in the index,
so older nano-coders can still read the logs. Deleting index.jsonl loses them; a new run then
asks again on the session's next turn.
Flags: --login github-copilot, --list-models PROVIDER, --trajectory SESSION_ID [--json|--markdown], --acp, --model provider/model (or AGENTIC_HARNESS_MODEL), --resume [SESSION_ID|last], --list-sessions [--all] [--json],
--config PATH, --verbosity LEVEL (-v), --sandbox off|workspace|read-only (or NANO_CODER_SANDBOX),
--allow RULE and --deny RULE (repeatable; added to the config's rules), --version (-V).
Before opening a PR, run cargo fmt (style in rustfmt.toml), cargo clippy --all-targets -- -D warnings
and cargo test; CI checks all three. To keep git blame past the one-time reformat, run
git config blame.ignoreRevsFile .git-blame-ignore-revs once.
Configuration
Create ~/.config/nano-coder/config.toml (every field is optional). Directories from before the rename (agentic-harness, for config and for data such as sessions) are moved to nano-coder at startup when the new ones don't exist yet. A symlink is left at the old path so an older nano-coder still finds them. If that compatibility symlink can't be created, a warning is printed and the link is retried on later starts; if the move itself fails, the old directory is used:
model = "anthropic/claude-sonnet-4-5" # provider/model
default_provider = "mock" # used when the model has no known provider prefix
temperature = 0.7 # or "default" to send none (see Temperature below)
max_tokens = 4096
max_iterations = 0 # LLM calls per user input (0 = unbounded)
system_prompt = "You are a helpful assistant with access to tools."
bash_timeout_secs = 600
persist_sessions = true
# session_dir = "/path/to/sessions" # default: <platform data dir>/nano-coder/sessions
session_titles = false # ask the model for a few-word title per session (--resume picker)
# title_model = "openai/gpt-4o-mini" # model for titles (default: the session's model)
auto_compact = true # summarize automatically when the context fills up
auto_compact_threshold = 0.8 # fraction of the context window
compaction_mode = "standard" # standard | smart (experimental, see Smart compaction)
# context_window = 128000 # override the window (providers can set it too)
verbosity = "normal" # quiet | normal | verbose | debug (or --verbosity)
renderer = "frame" # frame (default: app-owned redraw on resize) | legacy
timestamps = true # prefix CLI messages with the local time (HH:MM:SS)
project_instructions = true # load AGENTS.md etc. (see Project Instructions)
project_instruction_files = ["AGENTS.md", "CLAUDE.md", ".claude/CLAUDE.md", ".github/copilot-instructions.md"]
user_instruction_files = ["~/.claude/CLAUDE.md", "~/.agents/AGENTS.md"] # yours, for every project
user_rules_dirs = ["~/.claude/rules"]
instruction_imports_outside_project = false # let project @imports / rule links leave the repo
plan_tools = true # offer the plan_* tools (see Task Plans)
outcome_tool = true # offer report_outcome (see Outcomes)
reminders = true # append <system-reminder> notes to tool results
memory = "on" # on | read_only | off — cross-session memory (see Memory)
# memory_dir = "/path/to/memory" # default: <platform data dir>/nano-coder/memory
memory_expiry_days = 90 # expire memories unused this long (0 = never)
[skills] # see Skills
enabled = true
dirs = [".agents/skills", ".github/skills", ".claude/skills"] # relative to the git root
user_dirs = ["~/.agents/skills"]
ai_lock = true # load skills pinned in ai.lock
fetch = true # fetch ai.lock commits missing from the spm store
allowed_hosts = ["github.com"] # hosts ai.lock entries may be fetched from ("*" = any)
[permissions] # see Permissions and Sandbox
builtin_rules = true # block destructive commands unless allowed
allow = [] # e.g. ["Bash(sqlite3 test.db *)"]
deny = [] # e.g. ["Bash(git push:*)", "Edit(**/.env)"]
protected_branches = ["main", "master", "trunk", "develop"]
[sandbox]
mode = "off" # off | workspace | read-only (or --sandbox)
writable = [] # extra writable paths, e.g. ["~/.local/state/myapp"]
network = true # false blocks outbound connections
tool_caches = true # workspace mode: allow ~/.cargo/registry, ~/.npm, ~/.cache, ...The default model is gpt-4o-mini on the mock provider, so the harness still works offline.
Providers
A model is written as provider/model. The first path segment picks the provider if it
names one; otherwise the whole string is a model on default_provider. So
openrouter/anthropic/claude-sonnet-4.5 sends anthropic/claude-sonnet-4.5 to OpenRouter.
Built-in presets:
| Provider | Kind | Base URL | API key env |
|---|---|---|---|
| openai | openai | https://api.openai.com/v1 | OPENAI_API_KEY |
| anthropic | anthropic | https://api.anthropic.com/v1 | ANTHROPIC_API_KEY |
| openrouter | openai | https://openrouter.ai/api/v1 | OPENROUTER_API_KEY |
| fireworks | openai | https://api.fireworks.ai/inference/v1 | FIREWORKS_API_KEY |
| groq | openai | https://api.groq.com/openai/v1 | GROQ_API_KEY |
| together | openai | https://api.together.xyz/v1 | TOGETHER_API_KEY |
| deepseek | openai | https://api.deepseek.com/v1 | DEEPSEEK_API_KEY |
| kimi | openai | https://api.moonshot.ai/v1 | MOONSHOT_API_KEY |
| mistral | openai | https://api.mistral.ai/v1 | MISTRAL_API_KEY |
| gemini | openai | https://generativelanguage.googleapis.com/v1beta/openai | GEMINI_API_KEY |
| qwen | openai | Model Studio (default: Standard, Singapore — all plans/regions below) | DASHSCOPE_API_KEY / BAILIAN_*_PLAN_API_KEY |
| ollama | openai | http://localhost:11434/v1 | — |
| llamacpp | openai | http://localhost:8080/v1 | — |
| github-copilot | github-copilot | from session token | GITHUB_COPILOT_OAUTH_TOKEN or --login (unofficial, see below) |
| mock | mock | — | — |
kind = "openai" means OpenAI Chat Completions, which also covers vLLM, LM Studio,
llama.cpp, DwarfStar ds4 and similar servers. kind = "anthropic" is the Anthropic
Messages API.
A [providers.<name>] table can add a new endpoint or override any field of a preset:
[providers.ollama] # point the preset at another host
base_url = "http://merlin.local:11434/v1"
default_model = "qwen3:8b" # used by `--model ollama`
[providers.ds4] # a custom OpenAI-compatible endpoint
kind = "openai"
base_url = "http://localhost:8100/v1"
extra_body = { think = false } # merged into every request body
[providers.openai]
drop_params = ["temperature"] # for models that reject temperature
[providers.openrouter]
headers = { "HTTP-Referer" = "https://example.com", "X-Title" = "nano-coder" }
extra_body = { provider = { sort = "throughput" } }
[providers.work]
kind = "anthropic"
base_url = "https://llm-gateway.example.com/anthropic/v1"
api_key_env = "WORK_GATEWAY_KEY" # or api_key = "..." (prefer the env var)
# api_key_command = "op read op://vault/gateway/key" # used when the env var is unset
timeout_secs = 300 # idle timeout: max silence between streamed bytes (default 600)
max_retries = 3qwen is Qwen Cloud (Alibaba Cloud Model Studio), e.g. --model qwen/qwen3.8-max.
Model Studio serves the same models through three plans, each with its own hostnames and
API-key variable. All of them speak OpenAI Chat Completions, so each is just a base_url
api_key_env:
| Plan | Region | Base URL | API key env |
|---|---|---|---|
| Standard API key | Singapore (International) | https://dashscope-intl.aliyuncs.com/compatible-mode/v1 | DASHSCOPE_API_KEY |
| Standard API key | China (Beijing) | https://dashscope.aliyuncs.com/compatible-mode/v1 | DASHSCOPE_API_KEY |
| Standard API key | US (Virginia) | https://dashscope-us.aliyuncs.com/compatible-mode/v1 | DASHSCOPE_API_KEY |
| Standard API key | China (Hong Kong) | https://cn-hongkong.dashscope.aliyuncs.com/compatible-mode/v1 | DASHSCOPE_API_KEY |
| Token Plan | Singapore (International) | https://token-plan.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1 | BAILIAN_TOKEN_PLAN_API_KEY |
| Token Plan | China (Beijing) | https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1 | BAILIAN_TOKEN_PLAN_API_KEY |
| Coding Plan | Singapore (International) | https://coding-intl.dashscope.aliyuncs.com/v1 | BAILIAN_CODING_PLAN_API_KEY |
| Coding Plan | China (Beijing) | https://coding.dashscope.aliyuncs.com/v1 | BAILIAN_CODING_PLAN_API_KEY |
The qwen preset defaults to Standard API key, Singapore (the first row). In
/settings → Add or edit a provider → qwen these eight endpoints are offered as a
Qwen / Model Studio endpoint picker (plus a custom URL), so you can switch plan and
region without retyping a URL. Picking a plan also points the provider at that plan's
API-key variable — the preset's DASHSCOPE_API_KEY becomes BAILIAN_TOKEN_PLAN_API_KEY
for Token Plan, and BAILIAN_CODING_PLAN_API_KEY for Coding Plan. Set the variable (or a
key command / literal key) at the key prompt that follows. API keys are also bound to a
region, so a key issued for another region still needs the matching endpoint row.
To configure an endpoint in the config file directly, override base_url and
api_key_env (append /compatible-mode/v1 to a workspace domain, e.g.
https://<WorkspaceId>.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1):
[providers.qwen] # Token Plan, China
base_url = "https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1"
api_key_env = "BAILIAN_TOKEN_PLAN_API_KEY"kimi is the Kimi API from platform.kimi.ai, e.g. --model kimi/kimi-k3 or
kimi/kimi-k2.7-code. The preset drops temperature (K3 fixes it) and sets
replay_reasoning = true, which sends each assistant message's reasoning_content back
as thinking models like K3 require. Set extra_body = { reasoning_effort = "low" } to
make K3 think less.
Temperature
temperature is a number, or "default" to send none so the model uses its own default.
It can be set for all models, for one provider, or for one model; the most specific
setting wins:
temperature = 0.7 # all models
[providers.groq]
temperature = "default" # every groq model uses its own default
[providers.anthropic.models."claude-sonnet-4-5"]
temperature = 0.3 # this model onlySome models accept no temperature, and for them the model default is the only option:
GitHub Copilot's reasoning models (GPT-5 and later, Grok, …) and any provider with
drop_params = ["temperature"] (such as the kimi preset). A number set for such a
provider or model is ignored with a warning at startup and when you switch to it; the
top-level temperature just doesn't apply to them. Anthropic accepts 0 to 1, so a higher
value is sent as 1, with a warning. /context shows the temperature in use and where it
comes from, and /settings edits it for the current model, its provider, or all models.
Thinking
thinking sets how much the model reasons before it answers: "default" sends nothing
(the model decides), "off" turns thinking off, and a level such as "low", "medium",
"high", "xhigh" or "max" asks for that much. Like temperature, it can be set for all
models, one provider or one model, and /thinking LEVEL overrides them for the session:
thinking = "medium" # all models that support it
[providers.anthropic.models."claude-opus-4-7"]
thinking = "xhigh" # this model onlyA model only gets levels it supports. They come from, in order: thinking_levels on the
model or provider; what the endpoint reports for the model; and a built-in table
of Claude (3.7 Sonnet and later) and OpenAI reasoning models (GPT-5 and later, o1/o3/o4).
Endpoints that report levels:
- GitHub Copilot:
/modelslists each model's levels, read with the same request as its context window. - Ollama: a model with the
thinkingcapability (/api/show) getsoff,low,mediumandhigh. Ollama refuses a level for a model without it, so none is sent then. - llama.cpp: the loaded model's chat template (
/props) decides. A template with anenable_thinkingswitch (Qwen 3 and the like) getsoffandon, and any named level turns thinking on; a template that takesreasoning_effort(gpt-oss) getslow,mediumandhigh.
Ollama and llama.cpp are recognized however the provider is named, the same way as for the context window. For other models, list the levels yourself:
[providers.together.models."deepseek-r1"]
thinking_levels = ["low", "medium", "high"] # add "off" if thinking can be turned offA level the model lacks is moved to the nearest one it has (the highest below it, else the
lowest), and a model with no known levels is sent nothing. A setting for the provider, the
model or the session warns when it is adjusted or ignored; the top-level thinking doesn't
warn, since it applies to every model.
How the level is sent depends on the API: reasoning_effort (Chat Completions, Ollama
included; Ollama ignores its native think flag on /v1), chat_template_kwargs for
llama.cpp (enable_thinking, or reasoning_effort for templates that take a level; it
ignores the top-level reasoning_effort),
reasoning.effort (Responses), and for Anthropic Messages adaptive thinking with
output_config.effort, or on Claude 3.7 to 4.5 a fixed budget_tokens (1024 for minimal, then
2048, 8192, 16384, 32768, 65536 up to max). The budget has to stay below max_tokens, so raise
max_tokens to use a large one. Anthropic takes no custom temperature while thinking, so
none is sent then. A matching key in the provider's extra_body (reasoning_effort, reasoning,
think, chat_template_kwargs, thinking, output_config) is sent instead, with a warning; the
configured level is then reported as overridden (not as sent), since the override's value is what
reaches the wire. The status bar shows
think LEVEL while a level is sent, /context shows it with where it comes from,
/settings sets it for the current model, its provider or all models (picking from the
levels the model supports), and the session log records it for each reply.
Vision
read_file returns images (PNG, JPEG, GIF and WebP, recognised by magic bytes) to
models that can view them: the result is a short text part (image/png, 1600×900, 131 KB)
plus an image attachment the model sees directly. Images over the provider's limits are
downscaled first (longest side 1568 px, under the model's byte cap), and re-encoded to
JPEG/PNG when the model doesn't accept the source type. Other binary files still return
the "looks like a binary file" error, which notes whether the current model supports
images.
Whether a model can see images comes from, in order: a vision = true|false override
(global, [providers.<name>], or [providers.<name>.models."<model>"], like thinking);
what the endpoint reports (GitHub Copilot /models capabilities.supports.vision and its
limits.vision, Ollama /api/show's vision capability, llama.cpp /props
modalities.vision); and a built-in assumption for current Anthropic and OpenAI model
families. A model that can't see images gets the text error with a hint to switch models
or set vision = true:
vision = true # all models
[providers.ollama.models."my-clip-model"]
vision = true # this model onlyA request carries only the newest few images the model allows (GitHub Copilot's
max_prompt_images, default 1); older ones become [image omitted: path (sent earlier)].
Each image counts toward the context estimate at a fixed cost from its size, so the status
bar and auto-compaction stay honest, and compaction replaces images with the same
placeholder. Images are stored once per session under <session>.attachments/ (named by
content hash), so session logs stay small and --resume still works; a missing file on
resume is sent as a text placeholder.
Other per-provider fields: replay_reasoning, max_tokens_param (max_tokens, or max_completion_tokens
which is the openai default), retry_initial_backoff_ms, retry_max_backoff_ms and
retryable_statuses.
The old top-level api_key / base_url still work. They apply to default_provider,
which becomes openai if it was mock.
List a provider's models with --list-models <provider> (OpenAI-compatible endpoints and
github-copilot).
Local servers on another machine (macOS)
macOS Local Network privacy can block nano-coder from reaching a model server on your
LAN (e.g. http://192.168.0.141:8888/v1 or http://merlin.local:11434/v1). Requests fail
with tcp connect error … No route to host (os error 65), while curl to the same URL
works. localhost is not affected.
- Allow it in System Settings → Privacy & Security → Local Network. The prompt and the entry go to the app that launched nano-coder (your terminal, or the editor running it over ACP). A locally built binary is unsigned, so macOS may treat each rebuild as a new app and ask again.
- Or forward a local port, which needs no permission:
ssh -N -L 18888:localhost:8888 [email protected](orsocat TCP-LISTEN:18888,bind=127.0.0.1,fork TCP:192.168.0.141:8888), then setbase_url = "http://127.0.0.1:18888/v1".
GitHub Copilot (unofficial)
The github-copilot provider uses a GitHub Copilot subscription by authenticating as the
VS Code Copilot Chat extension: a device-flow login with VS Code's OAuth client ID, an
exchange for a short-lived Copilot session token, and model calls with VS Code's editor
headers. Each model is routed to the upstream API Copilot serves it through — Chat
Completions by default, the OpenAI Responses endpoint for gpt-*/grok-*/oswe*/mai-*
models (e.g. gpt-6-astra, which Copilot serves only via Responses), and the Anthropic
Messages endpoint for Claude 4.x/5.x. This is the approach several open-source agents
(e.g. pi) take, but it is not a GitHub-sanctioned integration. It may breach GitHub's
terms or your organisation's Copilot policy, it can break without notice, and misuse could
get an account flagged. It is never used unless you select it.
nano-coder --login github-copilot # interactive; saves the OAuth token (0600)
nano-coder --list-models github-copilot
nano-coder --model github-copilot/gpt-4.1Headless workers can't do the device flow; set GITHUB_COPILOT_OAUTH_TOKEN to a token from a
previous login instead (credentials live in <data dir>/nano-coder/github-copilot.json).
Tool follow-ups are sent with X-Initiator: agent, so a turn is billed like one VS Code
request. GITHUB_COPILOT_DOMAIN selects a GHE.com host.
The sanctioned route is the Copilot SDK, which drives the Copilot CLI's own agent loop rather than exposing the model.
Retries
Retry behaviour comes from unreal-agent's retry logic (MIT). Connection errors and
HTTP 408/409/425/429/5xx/529 are retried with exponential backoff (1s doubling to
30s, minus up to 20% jitter, 5 retries). Retry-After headers are honoured, and so is
"try again in Xs" in rate-limit messages. Overloads (overloaded_error,
server_is_overloaded, 529) back off from 10s up to 60s. Errors that can never succeed
on retry fail immediately: authentication and permission errors, invalid requests,
context_length_exceeded, quota and billing errors, and policy errors.
Output and Verbosity
In the interactive CLI, answers stream in as the model writes them. Output detail is set
with /verbosity, --verbosity, or verbosity in the config (default normal):
| Level | Shows |
|-------|-------|
| quiet | Final answers only |
| normal | Streamed answers, collapsed thinking, one line per tool call and its result |
| verbose | Also the first lines of each tool's output |
| debug | Also lifecycle hook events ([hook] ...) |
Thinking. Reasoning streams as a single line that updates in place
(∴ Thinking: ...) and becomes ∴ Thought for 3.1s · 812 chars when the answer starts.
Ctrl-O switches to showing thinking in full, including the block that is streaming. At
the prompt it prints the last thinking in full. Press it again to collapse. Reasoning is read
from reasoning_content / reasoning fields (DeepSeek, llama.cpp, vLLM, OpenRouter, Ollama),
from <think>...</think> in the content, and from Anthropic thinking blocks. Anthropic
thinking blocks are kept with the conversation so tool use keeps working when thinking is
enabled (e.g. extra_body = { thinking = { type = "enabled", budget_tokens = 4000 } }).
Tool calls show as ● bash ls -la, then ⎿ with the first line of the result (or the
error).
Input. On a terminal, input is read key by key. While a turn runs, what you type shows
on the status line; Enter sends it to steer the running turn, Ctrl-Enter (or Cmd-Enter)
adds it to the message queue (one queued message runs per following turn; /queue lists,
edits and removes them), and Esc Esc or Ctrl-C cancels the turn. The prompt supports
editing: Left/Right move the cursor, Home/End (or Ctrl-A/Ctrl-E) jump to the start/end,
Alt/Option-Left/Right (or Alt-B/Alt-F) move by word, and Up/Down recall submitted lines from
the session's input history (Down past the newest restores what you were typing). The mouse is never captured, so the
terminal keeps its native behaviour — the wheel scrolls the scrollback and drag selects text. Backspace and Delete remove the character before/under the cursor, Ctrl-U clears the
input, Ctrl-W deletes the word before the cursor, and Esc Esc (twice within a second)
clears the whole input at the prompt. At the prompt between turns,
Ctrl-Enter (or Cmd-Enter) inserts a newline without sending, and pasted text keeps its line breaks as a single multi-line input
instead of sending line by line. Ctrl-D exits on an empty line.
Streaming uses server-sent events. Set stream = false on a provider whose endpoint doesn't
support it. ACP mode doesn't stream text, but sends each response's reasoning as an
agent_thought_chunk update.
Every message (your prompt, answers, tool calls and results, thinking, notes) starts with
the local time as HH:MM:SS. The prompt's time is rewritten when you press Enter, so it
shows when the message was sent. Turn this off with timestamps = false.
Renderer. By default (renderer = "frame") nano-coder uses an app-owned frame
renderer: it composes the whole screen — transcript, input
editor, then the status bar as the last line — into one frame and diff-renders it through a
single writer, wrapped in synchronized-output markers so a half-drawn frame is never visible.
On a width change it re-renders the entire frame (clearing the screen and scrollback) instead
of relying on the terminal's native reflow, and resizes are debounced (~40ms) so a drag
settles on one clean redraw. The / command menu is drawn under the input editor as part of
the frame. It never captures the mouse. Setting renderer = "legacy" (or picking it in
/settings) switches back to the scroll-region renderer, where the terminal owns the
scrollback and reflows history itself on a resize, with the status line pinned to the bottom.
Status Line and Compaction
In an interactive terminal the bottom row shows the provider/model, the working directory
(home shown as ~; on a narrow terminal the middle directories collapse to …, as in
~/…/src/providers, before other items are dropped), context usage
(~ marks an estimate; without it the figure is anchored to the provider's reported usage),
a fill bar, message count, session input/output tokens, the auto-compaction threshold and
count (labelled smart-compact when compactions are smart, else auto-compact), the active mode when it is plan or auto (the default normal is not shown, to
save space), and what the agent is doing. With a GitHub Copilot model it also shows the session's
AI Credits (0.4 AIC), summed from the total_nano_aiu each response reports. In the legacy
renderer it uses a terminal scroll region (DECSTBM); the default frame renderer instead composes
the bar as the frame's final row. Either way it follows resizes, and is off when stdin/stdout isn't a TTY or AGENTIC_NO_STATUS is set. The conversation is kept
directly above the status line (empty space collects at the top), so shrinking the window
drops empty rows rather than pushing the conversation out of view.
The context window comes from, in order: context_window in the config, context_window
on the provider, the window the endpoint reports, a built-in table of known models, then
128k. /context shows which one applied. If a provider rejects a request as too long, the
harness takes the limit from the error, compacts, and retries once.
The endpoint is asked at startup and on /model (at most 5 seconds; skipped when config
sets the window). The window a server has loaded is preferred over the model's maximum:
| Server | Source |
|---|---|
| vLLM | /v1/models max_model_len |
| DwarfStar ds4, OpenRouter, Together, Kimi | /models context_length |
| Groq / Mistral | /models context_window / max_context_length |
| llama.cpp | /props n_ctx (per slot) |
| LM Studio | /api/v0/models loaded_context_length |
| Ollama | /api/ps for a loaded model, else num_ctx; never the model maximum, since Ollama runs with a smaller default |
| GitHub Copilot | /models max_prompt_tokens |
Compaction asks the current model to summarize older messages, keeping the recent tail
(up to 20k tokens, never starting at a tool result). Auto-compaction runs before a model
call when usage passes the threshold, or earlier if the prompt would leave less than the
reserved output room (MIN_OUTPUT_RESERVE, plus an estimation margin) within the window.
It won't run again until the context has grown by
another 10% of the window, so a context that can't shrink isn't summarized on every call.
If summarizing fails, the older messages are dropped with a note. The session log records
the new conversation, so --resume continues from it.
Smart compaction (experimental)
A summary is lossy: whatever it leaves out is gone for the agent, even though the session
log still has every original message. Smart compaction keeps the whole history reachable.
It is off by default (compaction_mode = "standard"); try it on one compaction with
/compact --smart, or set compaction_mode = "smart" to use it for auto-compaction too.
It needs a session log, and falls back to a standard summary when persist_sessions is off.
- Message IDs. A message's ID is the line of the session log where it was first
recorded, shown as
#N. IDs are stable across compactions and--resume. - Recorded detail. Each assistant message in the log also keeps the model's reasoning
text (
thinking), the request's tokenusageand its wall-clockduration_ms. These are for inspecting a session later (see #57) andhistory_read, which includes the reasoning. These recorded log fields themselves are never sent back to the model. This is separate fromreplay_reasoning(see above): when a provider has reasoning replay enabled, its reasoning blocks are still returned to the model as that provider requires. - Citing summary. The summarizer sees each message labelled
[#N]and is asked to cite(#N)for details whose exact text may matter (errors, commands, outputs, the user's wording) instead of copying them. The summary ends with the range it covers and a note that the originals can be retrieved. Tool results clipped at compaction point at their ID. - History tools. Once a smart summary is in the context, the agent gets two tools
over the current session's log (and only that one):
history_search(pattern, role?, before?, after?, order?, limit?)- case-insensitive regex (or plain text) search. Newest first by default (order=oldestfor early history); each matching message is one#N role name (time): snippetline with up to three snippets (and a[K matches]count when more matched). Earlier history lookups are not searched.history_read(id, max_output_length?)- one message in full, bounded like other tool output (the whole is spilled to a file when it is longer). Before a smart compaction the tools are not offered, so they cost nothing.
To judge whether it helps, the log records each compaction's mode, model and the
summarized line range (on the replace record), and each turn's history_calls (on
turn_end). /context shows the mode and this session's history-tool use.
scripts/compaction-report.py [SESSION_DIR] tabulates compacted sessions by model and mode:
compactions, turns after the first compaction, turns that used the history tools, and
reported outcomes. eval/compaction/ has a harness that compares the two modes across
providers and models on synthetic or forked real sessions (see its README). See #29 for the design.
Project Instructions
nano-coder reads the instruction files you already have for other tools, including Claude Code's, so there is nothing to duplicate.
Your files. Each file in user_instruction_files that exists (default ~/.claude/CLAUDE.md
and ~/.agents/AGENTS.md) is loaded first, under a "Your instructions" heading, followed by
the rules in user_rules_dirs (default ~/.claude/rules).
The repository's files. When a session starts, the harness looks in every directory from
the git root (the nearest ancestor containing .git) down to the working directory. Outside
a repository only the working directory is checked. In each directory:
- the first file found from
project_instruction_filesis used, soAGENTS.mdwins overCLAUDE.md, then.claude/CLAUDE.md, then.github/copilot-instructions.md; CLAUDE.local.md(personal, not committed) is loaded as well, after it.
Rules in .claude/rules/**/*.md at the git root are loaded after the root directory's
files. These files are appended to the system prompt, root first, under a "Repository
instructions" heading that tells the model to follow them. So the model has them before it
makes any change, without having to decide to read them. Block-level <!-- ... --> comments
are removed first. Each file is capped at 32 KiB and the total at 64 KiB; a file over the
limit is named, with a note to read it with read_file.
Imports. CLAUDE.md, CLAUDE.local.md, rules and imported files can pull in other files
with @path, as in Claude Code: relative to the importing file, ~/ allowed, up to four
levels deep, each file once. Code spans and fenced code blocks are skipped, \ escapes a
space, and a path that doesn't exist is treated as a mention. AGENTS.md has no import
syntax, so @ there is never expanded. Imports in repository files may not leave the
repository: a committed CLAUDE.md could otherwise send ~/.ssh/... to the model provider.
They are skipped with a warning in the banner and /context, and so are rules that are
symlinks to files outside it. Set instruction_imports_outside_project = true to allow them.
Your own files can import from anywhere.
Rules for some paths. A rule with paths front matter loads only when it is needed:
---
paths:
- "src/api/**/*.{ts,tsx}"
---
All API endpoints must validate their input.Patterns are relative to the git root (** crosses directories, * and ? don't, {a,b}
alternatives).
Deeper directories and path rules load lazily. Instruction files deeper than the working
directory, such as pkg/AGENTS.md, and rules whose paths match are attached to the result
of the first read_file, write_file or edit_file call that touches a matching path, once
each. After a compaction they are attached again the next time they apply. Files that the
bash tool touches don't trigger this.
Instructions are read again when a session is resumed, so edits take effect. /context lists
the loaded files, the rules waiting for a matching path, and anything skipped. To turn loading
off (yours and the repository's), set project_instructions = false, or set
AGENTIC_NO_PROJECT_INSTRUCTIONS.
Skills
A skill is a folder with a SKILL.md: YAML front matter with a name and a description,
then instructions, plus any scripts or reference files it needs. Only the name and description
of each skill go into the system prompt, under a "Skills" heading (8 KiB at most). When a task
matches one, the model calls load_skill, which returns the instructions (32 KiB at most),
the skill's directory, and a list of its other files for read_file.
Skills are found in this order, and the first skill with a given name wins:
- Repository:
.agents/skills,.github/skillsand.claude/skillsunder the git root, at any depth up to four folders (skills/<group>/<skill>/SKILL.mdworks). ai.lock: skills pinned by spm. Each locked skill is loaded, and so is each skill bundled in a locked plugin (its.claude-plugin/plugin.jsonskillsfolder, defaultskills/). Hooks, commands and MCP servers in plugins are not loaded.- User:
~/.agents/skills.
nano-coder reads ai.lock itself and never runs spm install, which edits the workspace
(.gitignore, vendor folders). Changes like that would end up in commits and PRs. Each
pinned commit is read from spm's store ($SPM_HOME/store, default ~/.spm/store) if it is
there. Otherwise it is fetched once into the nano-coder cache (<cache dir>/nano-coder/skills).
ai.lock is committed to the repository, so its entries are checked the way spm checks them:
a full 40-character commit, a store key that matches the URL and commit, and paths that stay
inside the checkout. Fetches are limited to skills.allowed_hosts ("file" allows
file://). With fetch = false, only commits already in the spm store or the cache are used.
An ai.json without an ai.lock is skipped with a warning, because unpinned references are
never resolved.
Skills are found again when a session starts or is resumed. Problems are listed at startup
and by /skills (and returned as _meta.skillWarnings over ACP); they never stop a session.
To turn skills off, set skills.enabled = false or set NANO_CODER_NO_SKILLS.
Task Plans
The plan_* tools give the agent a plan that lives outside the conversation. This helps most
with small context windows: the model can write down what it has done and learned, then let the
conversation be summarized without losing track.
plan_add- add steps (title strings or{title, after, note}), and optionally set thegoal. Items get numeric ids;afterlists items that must be finished first.plan_update- set an item'sstatus(pending,in_progress,done,blocked,dropped), rename it, or add
