mybuff
v1.5.0
Published
A terminal coding agent that can read, edit, and run your code.
Readme
mybuff
A small terminal coding agent. It talks to an OpenAI-compatible API, can read and edit your files, can run shell commands, and can search a codebase — asking for approval before anything changes.
Its architecture and terminal UI follow Freebuff: agent definitions pick the tool set, reads are batched, and the context is pruned before every request. See Architecture and Interface.
Install
npm install -g mybuffOr run it from a clone:
npm install
npm startUsage
cd your-project
mybuff # interactive session
mybuff "add a .gitignore" # run one prompt, then exit
mybuff --model qwen3-coder-next # override the model for this run
mybuff --agent lean # minimal tool set + prompt, lowest token cost per call
mybuff --resume # continue the last session in this directory
mybuff --no-stream # wait for the whole reply instead of streamingReplies stream token-by-token by default, so you see progress instead of
waiting for the full response. Set MYBUFF_STREAM=0 or pass --no-stream to
turn it off.
Works in any project directory; it uses your global config at
~/.mybuff/.env unless the folder overrides it.
Models
mybuff ships a curated set of two coding models:
| Model | Notes |
|---|---|
| deepseek-v4.1-flash | default · 1M context · fast and cheap |
| qwen3-coder-next | 256K context · code-focused |
/models lists them and /model <number|id> switches. Change the set in
src/config.ts (DEFAULT_CODING_MODELS), in MYBUFF_MODELS, or in mybuff.json:
MYBUFF_MODELS="deepseek-v4.1-flash,qwen3-coder-next" mybuff{ "codingModels": ["deepseek-v4.1-flash", "qwen3-coder-next"] }First run
If mybuff can't find an API key, it walks you through setup:
👋 Welcome to mybuff — let's get you set up.
Base URL [https://api.cheaperinference.com/v1]:
API key: verifying key… ok — 99 models available
Default coding model:
1. deepseek-v4.1-flash (default)
2. qwen3-coder-next
Choose [1]:
Saved to ~/.mybuff/.envThe key is written to ~/.mybuff/.env (masked as you type, existing lines
preserved) so one setup works in every project. Re-run the wizard any time by
deleting that file or the API_KEY line.
Configuration
Configuration is layered. Later sources win:
- Real environment variables
.envin the current directorymybuff.jsonin the current directory~/.mybuff/.env- Built-in defaults
.env:
API_KEY=your_provider_key
BASE_URL=https://api.cheaperinference.com/v1
MODEL=deepseek-v4.1-flash
MYBUFF_MODELS=deepseek-v4.1-flash,qwen3-coder-next # optional
MYBUFF_AD=try my other tool at example.com # optional ad line
MYBUFF_GRANT=write_workspace,run_commands # optional session grants
MYBUFF_KEEP_RESULTS=4 # recent tool results kept in full
MYBUFF_MAX_TOKENS=120000 # per-request ceiling; excess is clipped
MYBUFF_AGENT=base # base | lite
EXA_API_KEY=your_exa_key # optional; enables web_search
MYBUFF_USAGE=1 # print a token summary per run
MYBUFF_SESSIONS_DIR=~/.mybuff/sessions # optional; where sessions are stored
MYBUFF_RG=/path/to/rg # optional; ripgrep binary to use
NO_COLOR=1 # optional; plain output (FORCE_COLOR=1 forces it on)mybuff.json (non-secret defaults):
{ "model": "qwen3-coder-next", "baseUrl": "https://api.cheaperinference.com/v1", "ads": ["one", "two"] }Tokens
The conversation is resent on every request, so the cost of a turn is dominated by the tool results already in it. Three passes keep that under control before every send:
- Elision. Tool results older than the last
MYBUFF_KEEP_RESULTS(default 4) have their bodies replaced with[older tool output elided]. The messages themselves stay, so tool_call pairing is preserved. Short confirmations (under 400 chars) are left alone — they are cheap. - Budget fitting. If the payload still exceeds
MYBUFF_MAX_TOKENS(default 120,000 estimated tokens), the largest results are clipped, newest last, until it fits. The full results remain in memory for recovery. - Batch reads.
read_filestakes several paths per call, so a three-file read costs one round trip instead of three. Each round trip avoided also avoids re-sending the whole conversation.
File reads are clipped at 100 KB and 2,000 lines per file; command output and
search results at 8 KB. Set MYBUFF_USAGE=1 to print prompt/completion totals
plus how many results were elided or clipped.
Web tools
read_url needs no setup: it fetches a page, follows redirects, strips scripts
and styles, and returns readable text (with the title and meta description).
web_search uses Exa. Add a key to make it work — there are
free credits on signup:
EXA_API_KEY=your_key # or MYBUFF_SEARCH_KEY to overrideWithout a key the tool returns a clear "not configured" message instead of failing, and the agent falls back to the local codebase.
Sub-agents
spawn_agents delegates to read-only sub-agents, which run in parallel and
report back a summary instead of flooding the main conversation with raw output:
| Sub-agent | Use it for |
|---|---|
| file-picker | find the files relevant to a task |
| researcher | gather facts and cite them |
| reviewer | critique a change for correctness and edge cases |
They are read-only by construction — none carries a mutating tool, none can
spawn further sub-agents, and none can ask the user. Spawning prompts for the
delegate capability; pre-grant it with MYBUFF_GRANT=delegate (this is how a
scripted run delegates). spawn_agents is offered by the base agent only:
a one-shot run has no way to answer the approval prompt.
Delegation is a trade, not a free win. Each sub-agent runs its own model loop, so for a one-line lookup it costs more than doing the work directly — it pays off when the alternative is pulling a lot of raw output into your own context.
Interface
The CLI is styled after Freebuff's — a green accent, boxed panels, and a spinner while the model is working:
╭─ mybuff ─────────────────────────────────────────╮
│ mybuff v1.5.0 · terminal coding agent │
│ │
│ model deepseek-v4.1-flash │
│ agent base │
│ cwd D:\projects\thing │
│ │
│ tools read_files, str_replace, write_file, … │
╰──────────────────────────────────────────────────╯
⏺ code_search {"pattern":"handleAuth"}
⎿ src/auth.ts
mybuff › Auth is handled in src/auth.ts.It is plain ANSI rather than a TUI framework, so it works in any terminal and
adds no dependencies. Colour is disabled automatically when output is piped or
NO_COLOR is set, and forced with FORCE_COLOR=1.
Sessions
The conversation is saved to ~/.mybuff/sessions/<project>.json after every
turn, so quitting no longer discards it:
mybuff --resume # continue the last session in this directory
mybuff --no-save # don't write oneSessions are per project directory and store only conversation state — never
the API key. The stored copy is pruned with the same policy used for requests,
so large file reads don't grow the file without bound. /clear erases both the
in-memory conversation and the saved session.
Architecture
The structure follows Freebuff's:
- Agent definitions (
src/agents.ts): each agent is{ id, toolNames, systemPrompt }, and the runtime only ever offers the model that agent's tools.baseis the interactive agent;litewithholdsask_userandsuggest_followupsfor one-shot runs where no human is attached. Pick one with--agentorMYBUFF_AGENT. - Tool registry (
src/tools.ts): one entry per tool, and the request carries exactly the agent's tools — not a hardcoded list. - Auto-read-only batching: independent read-only tool calls in the same assistant message run in parallel; mutating calls stay ordered. Sub-agents spawned in one call also run in parallel.
- New tools:
read_url(dependency-free readable-text extraction),web_search(Exa REST viafetch— no SDK), andspawn_agents(nested read-only agents with a recursion guard). - Stable prompt prefix (
src/prompt.ts): the system prompt is assembled once so identical bytes lead every request — nothing is re-injected mid-conversation, which is what would break the provider's prompt cache.
Ads and updates
At the start of an interactive session mybuff may print one text ad and, if a newer release is on npm, a one-line update hint. Both are quiet by default:
- The ad slot is empty unless you set
MYBUFF_ADormybuff.json→ads(an array of strings). Multiple ads rotate once per day. - The update check asks the npm registry for the latest version. It times out after 2.5s, never blocks startup, and is skipped when running from a clone.
Slash commands
| Command | Effect |
|---|---|
| /help | show the command list |
| /models | list available models with price, context, and capabilities |
| /model [n\|id] | show or switch the model (by number or id) |
| /agent [id] | show the agent definitions, or pick one for the next launch |
| /status | session info and token usage |
| /todos | show the agent's current task list |
| /skills [name] | list skills, or preview one |
| /grants | list, revoke, or clear permissions granted this session |
| /clear | clear the conversation history |
| /exit | quit |
Skills
Skills are reusable instruction sets the agent loads on demand through the
skill tool. Add one by creating a folder with a SKILL.md inside:
.agents/skills/my-skill/SKILL.md---
name: my-skill
description: What this skill does (shown to the agent for discovery)
---
# My Skill
Instructions the agent follows when it loads this skill.mybuff ships built-in skills you get with no setup: debug-and-fix,
git-workflow, and code-review.
Discovery order is built-in, then ~/.agents/skills/ (global), then
.agents/skills/ (project) — later sources override earlier ones of the same
name. List skills with /skills and preview one with /skills <name>.
Tools
The agent can call these tools. Writing, editing, and command execution each require an approval prompt; every other answer than an explicit yes denies.
The eight core tools mirror Freebuff's base3 harness; the last three are CLI surface, offered only when a human is attached.
| Tool | Purpose |
|---|---|
| read_files | read several files in one call; supports { path, offset, limit } windows |
| str_replace | replace an exact, unique string in a file |
| write_file | create or overwrite a file |
| run_terminal_command | run a shell command in the working directory |
| code_search | ripgrep search with line numbers; falls back to a built-in, .gitignore-aware scanner handling -i, -g, -w, -F, -l, -c, -v, -t, --type-not, -m, -A/-B/-C |
| glob | find files by glob pattern |
| list_directory | list a directory's entries |
| write_todos | the agent's task list, shown with /todos |
| read_url | fetch a URL and return its readable text |
| web_search | search the web (Exa) for current information |
| skill | load a reusable skill's instructions on demand |
| ask_user | multiple-choice questions on non-obvious decisions |
| suggest_followups | suggest ~3 next steps at the end of a turn |
| spawn_agents | delegate to read-only sub-agents that report back |
The older singular names (read_file, edit_file, run_command) still work as
aliases, so existing skills and prompts keep functioning.
Permissions
Each tool maps to a capability, and mybuff asks before using one:
| Tool | Capability |
|---|---|
| read_files, code_search, glob, list_directory, skill | read_workspace |
| write_file, str_replace | write_workspace |
| run_terminal_command | run_commands |
| write_todos | agent_control |
| read_url, web_search | network |
| spawn_agents | delegate |
| ask_user, suggest_followups | human_in_loop |
Only mutating tools and spawn_agents prompt. Read-only tools never ask.
At the prompt, y allows the single action, a grants the capability for
the rest of the session (so repeat actions stop prompting), and anything else
denies. /grants lists what you've granted, /grants revoke <capability>
removes one, and /grants clear removes them all.
Pre-grant scopes for a whole session with MYBUFF_GRANT (comma-separated) — the
hook a sponsored run uses to grant permissions up front:
MYBUFF_GRANT=write_workspace,run_commands mybuffAvailable capabilities: read_workspace, write_workspace, run_commands,
network, human_in_loop, delegate, agent_control.
Tests
npm test # node --test via tsx; runs offline
npm run typecheck
MYBUFF_RG=/path/to/rg npm test # also exercise the ripgrep pathThe suite covers the tool registry and every tools interface, the context
pruning passes (elision, budget fitting, purity), agent/tool wiring, session
persistence, the interactive decisions, and the full code_search flag set —
run against both engines when a ripgrep binary is available.
read_url is exercised against a local HTTP server, so the suite runs offline,
except for the JS-rendering case, which skips when playwright is not installed.
Tests live in test/ and are dev-only: they are not part of the published
package.
Verification status
What has been checked, and what has not:
| Area | Status |
|---|---|
| Tool behaviour (all 14) | covered by npm test and tsc --noEmit |
| Context pruning / token estimator | covered by npm test |
| read_url | live over real HTTPS, plus a local-server test |
| spawn_agents | live end-to-end, nested sub-agent loop verified |
| web_search | endpoint, x-api-key auth, and error parsing verified live; a successful search is not (no key available) |
| code_search with ripgrep | verified — run against a local ripgrep 15.0.0 binary via MYBUFF_RG; all 17 flag tests pass on both engines |
| .gitignore handling | covered by npm test (root + nested rules, ! negation, **) |
| Session save / resume | covered by npm test, and live: a value set in one process was recalled by --resume in another |
| Terminal UI | verified with forced colour: escapes emitted, boxes column-aligned, plain output when piped |
| read_url JS rendering | verified — a client-rendered page returned its JS-injected text; the test skips where playwright is absent |
| Interactive decisions (y/N/a, EOF→deny, grants, answer parsing) | covered by npm test (extracted into src/interactive.ts) |
| Interactive rendering in a real TTY | unverified — no terminal is available in the build environment |
| Output quality | not measured — no A/B of answer quality has been run |
Known limitations:
Delegation costs more for small tasks. Each sub-agent runs its own model loop; a one-line lookup measured 5 calls / ~11k tokens. It pays off when the alternative is pulling a lot of raw output into your own context.
Per-request overhead is ~2.2k tokens (
base) of system prompt plus tool schemas, deliberately spent to avoid extra round trips. On a short question that is a net cost; on multi-step work it is a net saving. Use--agent lean(~0.9k) when you want the floor instead — see the table above.Without ripgrep,
code_searchis a plain recursive scan. It honours.gitignore(root and nested, including!negations and**) and the full documented flag set, but an unrecognised flag or-ttype is reported as ignored rather than silently accepted.The UI is ANSI-styled, not the real OpenTUI/React component tree Freebuff uses. It matches the look (accent colour, boxed panels, spinner, tool bullets) but not the interactive widgets — model picker panels, clickable cards, in-line diffs. Those need a TUI framework.
Agent-by-agent token overhead, measured on this machine:
| Agent | Tools | Prompt | Overhead per request | |---|---|---|---| |
base| 14 | 1932 chars | ≈ 2163 tokens | |lite| 11 | 1358 chars | ≈ 1612 tokens | |lean| 6 | 549 chars | ≈ 925 tokens |read_urldoes not run JavaScript unless you install a headless browser yourself:npm i playwright. Without it, client-rendered pages return a note explaining that, rather than a silent empty result.
License
ISC
