@justin06lee/yagami
v0.12.0
Published
Self-hosted Anthropic- and OpenAI-compatible API backed by your signed-in coding-agent CLIs, plus a zero-config library for driving them from your own apps — no API keys needed.
Maintainers
Readme
yagami
Your signed-in coding-agent CLIs as one self-hosted Anthropic- and OpenAI-compatible API. Claude Code, Codex, OpenCode, Gemini CLI and any ACP agent — point any Anthropic or OpenAI client at your own subscriptions, or embed the engine as a library with zero config.
yagami does the T3-Code trick, generalized: it drives the coding-agent CLIs you already installed and logged into — Claude Code through the Agent SDK (pathToClaudeCodeExecutable), Codex through codex exec, and OpenCode, Gemini CLI, Copilot, Cursor, Qwen Code, Kimi, Goose and the rest of the ACP registry through the Agent Client Protocol. No API keys from any vendor, no separate auth — each spawned engine uses the same login your terminal sessions do.
There are two doors in, and they're deliberately different:
| | Who it's for | What you get |
|---|---|---|
| Server — the yagami binary | A machine that should hand out an API key (a home server, a box your other apps talk to) | yagami start prints a URL + API key that works with any app that accepts an Anthropic or OpenAI base URL + key |
| Library — @justin06lee/yagami | An app running on the machine with the signed-in CLIs | new Yagami() — no URL, no API key, nothing to configure. It finds the CLIs itself and syncs with the binary's config |
Personal use only. This exists so you can point your own tools at your own subscriptions. Offering subscription-backed access to other people is against every one of these vendors' terms. Keep the endpoint private and don't share keys.
Server mode
make # bun install + build + install `yagami` onto your PATH (~/.local/bin)
yagami start # first run generates + saves an API key and prints itmake update stops any running yagami server, rebuilds, reinstalls, and restarts it. Or install from npm: bun add -g @justin06lee/yagami.
yagami v0.12.0
listening http://127.0.0.1:8787
provider claude — /Users/you/.local/bin/claude (2.1.238 (Claude Code))
also codex, opencode (use model "<provider>:<model>")
api key ygm_…
Connect apps — either dialect, same key (`yagami key` prints ready-to-paste env exports):
Anthropic apps baseURL http://127.0.0.1:8787 (ANTHROPIC_BASE_URL / ANTHROPIC_API_KEY)
OpenAI apps baseURL http://127.0.0.1:8787/v1 (OPENAI_BASE_URL / OPENAI_API_KEY)Apps can't be reached by a key alone — they need somewhere to route — so the pair is always URL + key. Most apps take them as a "base URL" field or the standard env vars; yagami key prints both ready to paste:
export ANTHROPIC_BASE_URL=http://127.0.0.1:8787 ANTHROPIC_API_KEY=ygm_...
export OPENAI_BASE_URL=http://127.0.0.1:8787/v1 OPENAI_API_KEY=ygm_...Anthropic-dialect apps speak to POST /v1/messages, OpenAI-dialect apps to POST /v1/chat/completions — same engine, same key, streaming included:
import Anthropic from "@anthropic-ai/sdk";
const anthropic = new Anthropic({ baseURL: "http://127.0.0.1:8787", apiKey: process.env.YAGAMI_KEY });
await anthropic.messages.create({ model: "sonnet", max_tokens: 1024, messages: [{ role: "user", content: "hello" }] });
import OpenAI from "openai";
const openai = new OpenAI({ baseURL: "http://127.0.0.1:8787/v1", apiKey: process.env.YAGAMI_KEY });
await openai.chat.completions.create({ model: "codex:gpt-5.6-sol", messages: [{ role: "user", content: "hello" }] });Or raw curl:
curl http://127.0.0.1:8787/v1/messages \
-H "x-api-key: ygm_..." -H "content-type: application/json" \
-d '{"model":"codex","max_tokens":64,"messages":[{"role":"user","content":"ping"}]}'
curl http://127.0.0.1:8787/v1/chat/completions \
-H "authorization: Bearer ygm_..." -H "content-type: application/json" \
-d '{"model":"codex","messages":[{"role":"user","content":"ping"}]}'Web search and fetch
/v1/messages accepts Anthropic's server tools, so a client can ask the engine to go look something up before it answers:
curl http://127.0.0.1:8787/v1/messages \
-H "x-api-key: ygm_..." -H "content-type: application/json" \
-d '{"model":"sonnet","max_tokens":600,
"tools":[{"type":"web_search_20260209","name":"web_search"}],
"messages":[{"role":"user","content":"What is the current stable Rust version?"}]}'Server tools run inside the engine — they map onto the CLI's own WebSearch and WebFetch, and the results are folded into the reply. Nothing changes about the endpoint's contract: it still never emits a tool_use block for you to execute. Every dated variant of a type works (web_search_20250305, web_search_20260209, …), tool_choice may be auto or none, and any other tool in the array is rejected — a custom tool would have to run on your side, and there is no round trip here to run it on. Only the claude provider serves them; asking codex or an ACP agent for them fails loudly rather than quietly answering without the lookup.
One caveat for OpenAI-dialect apps: the model field still routes through yagami's providers — set it to a model your CLIs actually serve (sonnet, codex:gpt-5.6-sol, opencode:…), not whatever gpt-* id the app defaults to.
MCP servers
/v1/messages also accepts Anthropic's MCP connector: declare the servers in mcp_servers, enable them with mcp_toolset entries in tools, and the engine connects to them for the turn. The model calls their tools as it works, and the activity comes back in the reply as mcp_tool_use / mcp_tool_result content blocks, exactly as the real API does it:
curl http://127.0.0.1:8787/v1/messages \
-H "x-api-key: ygm_..." -H "content-type: application/json" \
-d '{"model":"opus","max_tokens":1024,
"mcp_servers":[{"type":"url","url":"http://127.0.0.1:52011/mcp","name":"editor"}],
"tools":[{"type":"mcp_toolset","mcp_server_name":"editor"}],
"messages":[{"role":"user","content":"What is the root node of the open scene?"}]}'This is how a program gets tools without running a tool loop: the tools live wherever the MCP server lives (a game editor, a desktop app, a box on the LAN), the turn stays one request/response, and the server sees a normal MCP client. hitbox, the Godot fork with Claude built in, works this way: the editor is the MCP server, yagami is the brain. authorization_token becomes a bearer header, tool_configuration.allowed_tools and per-toolset configs narrow what the model may call, and mcp_tool_use/mcp_tool_result blocks echoed back in assistant history are accepted (they replay as a short trace when no cached session matches). Servers must speak streamable HTTP or SSE; the claude provider is the only one that connects them, and the turn cap rises from 1 to 64 model turns while they are connected.
Library mode
For apps that run on the machine with the signed-in CLIs — a desktop app, a script, anything embedding the engine in-process. No server, no URL, no API key: new Yagami() auto-detects the CLIs and reads ~/.config/yagami/config.json if the binary has one, so library and server stay in sync automatically.
import { Yagami } from "@justin06lee/yagami";
const yagami = new Yagami(); // that's it — hooks straight into the host's CLIs
// Anthropic SDK shape:
const msg = await yagami.messages.create({ messages: [{ role: "user", content: "hello" }] });
for await (const ev of yagami.messages.create({ messages: [...], stream: true })) { /* Anthropic stream events */ }
// OpenAI SDK shape, same engine:
const completion = await yagami.chat.completions.create({ messages: [{ role: "user", content: "hello" }] });
for await (const chunk of yagami.chat.completions.create({ messages: [...], stream: true })) { /* chat chunks */ }
// Models across every installed harness:
const { data } = await yagami.models.list();Options are for overrides only (new Yagami({ defaultModel: "sonnet" }), { defaultProvider: "codex" }, { syncHostConfig: false }, …); the zero-argument form is the intended use. For lower-level control the engine underneath is yagami.engine (a YagamiEngine — complete() returns cost/session/provider metadata, stream() returns raw SSE events), and you can hand-pick providers instead of auto-detecting:
import { YagamiEngine, ClaudeProvider, AcpProvider } from "@justin06lee/yagami";
const engine = new YagamiEngine({
providers: [new ClaudeProvider(), new AcpProvider({ id: "gemini", label: "Gemini", command: "gemini", args: ["--acp"] })],
});
const { response, costUsd } = await engine.complete({ messages: [{ role: "user", content: "hello" }] });Every provider implements one small Provider contract (run(turn) → normalized session/text/thinking/done events, plus listModels() and version()), so adding a harness that isn't ACP-capable is one file. Failures are typed: AuthRequiredError (carries the login command), ProviderNotInstalledError (carries the install hint), ProviderError. Every CLI yagami starts is ended with everything it started in turn — a launcher that re-executes itself, an npm shim, an agent's MCP servers — and nothing is left waiting on an agent that never answers: an ACP handshake has 30 seconds, a model-list or version probe 20, and a turn aborted before its session exists closes the agent. A consumer that stops reading a turn early ends its process too.
Building a UI on Claude Code
Yagami/YagamiEngine are completions-only by design (server tools aside, they never hand you a tool call to run). To build an actual coding UI — tools, permissions, plan mode, a warm session across turns — use AgentSession, which wraps the full Claude Code agent with the lifecycle the interactive terminal gives you for free:
import { AgentSession } from "@justin06lee/yagami";
const session = new AgentSession({
cwd: "/path/to/project",
parity: "terminal", // load your CLAUDE.md, skills, hooks, .mcp.json — like the CLI
appName: "my-app", // reported to Claude as the client
onPermission: async (req) => {
// Your approve/deny UI. Policy stays here; yagami owns the state machine.
const ok = await showDialog(req.toolName, req.input);
return ok ? { behavior: "allow" } : { behavior: "deny", message: "user declined" };
},
});
session.send("fix the failing test"); // process starts here and stays warm
for await (const msg of session) { // raw SDKMessages — render however you like
render(msg);
if (msg.type === "result") break;
}
session.send("now add a test for the edge case"); // next turn resumes the same session
await session.interrupt(); // the CLI's Esc
await session.setModel("opus"); // the CLI's /model
await session.setPermissionMode("plan"); // shift+tab
// the CLI's /rewind: restore tracked files to their state at a user
// message's uuid (needs options: { enableFileCheckpointing: true })
await session.rewindFiles(userMessageUuid);
session.close();This resolves the parts of embedding Claude Code that every host would otherwise reimplement identically — process lifecycle, session resume, interrupt, settings parity, and the permission state machine. What stays yours are the genuinely app-specific choices: rendering the SDKMessage stream, deciding what to auto-approve, and picking the working directory. parity is "terminal" (load user+project+local settings, matching your CLI), "project" (project+local only), or "isolated" (load nothing — reproducible, no personal config). The permission fallback defaults to "deny", so a session is safe before the UI is wired up; autoAllow/autoDeny skip the handler for named tools.
For a lower-level handle, claudeCodeSession(prompt, { options }) returns the raw Agent SDK Query.
Building a UI on any other harness
The same idea works for the non-Claude harnesses — verbatim. Codex and every ACP agent implement SessionProvider.openSession(): a live, warm session on the harness's own engine (codex app-server — what the Codex TUI runs on; a persistent ACP connection for OpenCode, Gemini, and friends), with the harness's own config, sandbox, and approval flow. Nothing is overridden unless you pass native overrides (or parity: "isolated", which switches the user's own MCP servers, plugins and apps off for a Codex thread); approval requests are forwarded to your handler exactly as the harness's own UI would prompt.
import { createProvider, isSessionProvider } from "@justin06lee/yagami";
const codex = createProvider("codex", {}, { appName: "my-app" });
if (isSessionProvider(codex)) {
const session = codex.openSession({
cwd: "/path/to/project",
permissions: {
decide: async (req) => (await showDialog(req.tool, req.input)) ? "allow" : "deny",
}, // "allow_always" answers like the TUI's "don't ask again"
input: {
// Codex request_user_input and MCP/ACP form or URL elicitations all
// arrive in this provider-neutral shape. Throwing safely cancels it.
respond: async (request) => renderInput(request),
},
});
for await (const ev of session.send("fix the failing test")) {
// normalized AgentEvents: session / turn / text / thinking / tool_call
// (started→completed, including Codex multi-agent operations) / permission
// / plan / done. A Codex subagent's own text and tool calls arrive too,
// tagged `thread` with the id of the thread its spawn_agent call started
}
session.send("now add a test"); // same warm thread, context carries
await session.interrupt();
await session.close(); // session.id resumes it later via { resume }
}ProviderSessionOptions takes cwd, model, resume, effort, systemPrompt (extra developer instructions where the harness supports them), permissions, optional input, mcpServers, and a native escape hatch (Codex: { sandbox, approvalPolicy, config }; ACP: { mode }). mcpServers ({ name: { url, headers?, allowedTools? } }, streamable HTTP) connects the session to your own tool servers on top of whatever the harness loads from its config. Codex asks before running their tools, and that ask reaches permissions.decide with server set to the MCP server's name, so a host can allow exactly its own tools. ACP agents take the servers only if they advertise HTTP MCP support, and they don't enforce allowedTools. A session provider reports sessionCapabilities.fork; when true, { resume, fork: true } branches at the tip and { resume, forkAt: turnId } branches through an exact turn event without mutating the source conversation. Input fields preserve labels, options, required/secret flags, primitive constraints, and URLs; omitting the handler declines safely instead of hanging a turn. ACP sessions also map effort onto the agent's thought_level option when it exposes one. The completion-turn run() path stays for API-style callers; sessions are for hosts that want the real interactive agent.
The server is also embeddable: import { startYagami } from "@justin06lee/yagami/server" — it takes the same fields as the config file plus log and providerInstances (hand-picked Provider objects instead of auto-detection).
Providers
A bare model id goes to the default provider (Claude Code unless you change it). "<provider>:<model>" routes to another harness; a bare provider id ("codex") means that harness's own default model. GET /v1/models and yagami models list everything that's actually installed, with ids ready to paste.
| Provider | Driven through | Resume | Images | System prompt | Thinking / effort |
|---|---|---|---|---|---|
| claude — Claude Code | Agent SDK → your claude binary | yes, forking | yes (+ documents) | native | native |
| codex — Codex CLI | codex exec --json (read-only sandbox) | yes | yes | emulated | effort only |
| opencode, gemini, copilot, cursor, qwen, goose, kimi, kilo, cline, auggie, amp, grok, droid, codex-acp, claude-acp | Agent Client Protocol over stdio | if the agent supports it | if the agent supports it | emulated | thought_level when exposed |
"Emulated" means the system prompt is folded into the user turn as a <system> block; unsupported thinking/effort are accepted and reported in x-yagami-ignored rather than rejected. Without native forking, a resumed session is single-use: a sibling branch of the same conversation falls back to transcript replay instead of corrupting the shared session.
Any other ACP agent works too — add it to config with its launch command:
{
"defaultProvider": "claude",
"providers": {
"codex": { "sandbox": "read-only" }, // "workspace-write" if you trust the callers
"gemini": { "path": "/opt/homebrew/bin/gemini" },
"my-agent": { "command": "my-agent", "args": ["acp"], "label": "My Agent" },
"goose": { "enabled": false } // hide a preset even if installed
}
}yagami doctor shows every known harness, whether it's installed and signed in, its version, and — for Claude — whether the bundled Agent SDK build matches your binary. Launch commands come from the ACP registry; sign-in hints are best-effort.
CLI
| Command | What it does |
|---|---|
| yagami start | Start the server (-p port, -H host, --provider <id> default provider, --claude <path>, --cors). Add --daemon to run it in the background (--log <file> overrides the default log at ~/.config/yagami/yagami.log) |
| yagami stop | Stop the running server: SIGTERM, then SIGKILL after 5 s if it's wedged. Only ever signals the process that wrote the state file — a pid recycled after a reboot or crash is left alone |
| yagami status | Show whether it's running, plus its providers, uptime, request count, and cumulative would-be API cost |
| yagami key | Print the URL + API key, plus ready-to-paste ANTHROPIC_*/OPENAI_* env exports for client apps |
| yagami models | List models across every installed provider (--provider <id> to filter) |
| yagami keygen | Generate another API key and save it to the config |
| yagami doctor | Check every harness CLI; --live sends one tiny real completion (--provider <id> to pick which) |
Every request is logged as one line (time, status, model, duration, cost, session) to stdout — or to the log file in --daemon mode.
Config
~/.config/yagami/config.json (override dir with YAGAMI_CONFIG_DIR):
{
"host": "127.0.0.1", // keep loopback unless you know what you're doing
"port": 8787,
"apiKeys": ["ygm_..."],
"defaultProvider": "claude", // provider for bare model ids
"defaultModel": "sonnet", // used when a request omits `model` (may be "provider:model")
"providers": { ... }, // see above; every preset is auto-detected when omitted
"cors": false
}Env overrides: YAGAMI_HOST, YAGAMI_PORT, YAGAMI_API_KEY, YAGAMI_PROVIDER, YAGAMI_DEFAULT_MODEL, YAGAMI_CLAUDE_PATH, YAGAMI_CODEX_PATH. The older claudePath / claudeConfigDir keys still work as shorthands for providers.claude. Library mode reads the same file (minus the server-only fields — host, port, keys), which is what keeps an embedded Yagami and the binary in agreement.
A config file that exists but isn't valid JSON is an error, not "no config": the binary refuses to start (it would otherwise save defaults plus a fresh key over your settings), and library mode warns and proceeds with auto-detection alone. Everything yagami survives on purpose — a failed model probe, a session cache it couldn't save, a host handler that threw — is reported on stderr; set YAGAMI_DEBUG=1 to also see the expected, recoverable kind, and in library mode setLogSink(fn) routes all of it wherever your app logs (setLogSink(() => {}) silences it, setLogSink(null) restores stderr).
How it works
- Engine: each request becomes one sandboxed turn on the chosen harness. Claude runs with
tools: [],settingSources: [](your CLAUDE.md/skills never leak into API completions),maxTurns: 1and a deny-all permission callback — a request with server tools or MCP servers lifts the turn cap (24 and 64) and allows exactly those tools; Codex runs in its read-only sandbox with no approvals (the prompt reaches it on stdin, never as an argument, so a message that starts with-is a message and a long replayed transcript can't overflow the argument list); ACP agents are moved to a plan/read-only mode when they offer one and every permission request is refused. All of them work in a throwaway directory. The API is text-in/text-out; a leaked key can burn tokens but never edit anything on the host — though note that agents other than Claude keep their own read-only tools, so they can still look at that empty directory. - Dialects:
POST /v1/messagesis native.POST /v1/chat/completionstranslates OpenAI shapes at the edge — system/developer messages fold intosystem,image_urlparts become image blocks, streams are re-emitted aschat.completion.chunkevents ending in[DONE], and thinking output rides along asreasoning_content. Errors on that path come back OpenAI-shaped too.GET /v1/modelsserves one merged shape both SDKs parse. - Multi-turn: the Messages API is stateless but harness sessions aren't. yagami hashes each conversation prefix (per provider) and remembers which session produced it; a follow-up request resumes that session and sends only the new user message. Unmatched histories fall back to replaying the transcript in a single prompt, and if a cached session turns out to be gone, the stale mapping is dropped and the request transparently retries via replay. The cache persists across restarts at
~/.config/yagami/sessions.json. - Streaming: every harness's output is normalized into deltas and re-emitted as a proper Anthropic SSE sequence —
message_start→ thinking/text content blocks →message_delta→message_stop(or the OpenAI chunk sequence on the chat-completions path). Claude and ACP agents stream tokens; Codex streams per message part. - Models:
GET /v1/modelsasks each installed CLI what it supports (Claude via the SDK, Codex via its app-server protocol, ACP agents via their session config) — probed once per process, then cached. Library callers also receive native model metadata when reported: reasoning levels/default, input modalities, fast/auto/adaptive-thinking flags, personality and multi-agent support, service tiers, and the provider's default model. Failed probes are skipped and retried next time; a static fallback list is served only if nothing answers (x-yagami-models-sourcesays which). - Auth:
x-api-keyorAuthorization: Bearer, compared in constant time. Binds to127.0.0.1by default and warns loudly on anything else. Request bodies are capped at 32 MB (the real API's limit;413beyond it), and a port that's already taken failsyagami startwith one line instead of a crash.
Extra response headers: x-yagami-provider, x-yagami-cost-usd (what the turn would have cost at API prices, when the harness reports it), x-yagami-session, x-yagami-ignored (accepted-but-unsupported params). /healthz answers { ok, service, version } to anyone (liveness); with a valid key it also reports the default provider, installed providers, the binary's path, uptime, request count, and the cumulative would-be cost — yagami status shows the same.
Limitations
- No client tool loop: custom
tools/tool_choice/ function calling are rejected with 400 (by design, see above), as aretool_use/tool_resultcontent blocks and OpenAItool/functionmessages. Tools reach the model only as Anthropic server tools or through the MCP connector, where they run on your MCP server inside the turn. - User messages may contain
text,image, anddocumentblocks (documents: Claude only; images: base64 sources only outside Claude);systemand assistant messages are text-only. Thinking blocks echoed back in assistant history are dropped, not rejected. A conversation whose history contains images/documents can only be continued while the server that produced it still has that session cached. - Assistant prefill (a trailing
assistantmessage) is emulated: the engine is instructed to continue from the prefill text, and the response carries only the continuation, like the real API. An accidentally repeated prefill is stripped from the reply, including mid-stream. max_tokens,temperature,top_p,top_k,stop_sequences(and their OpenAI counterparts, pluspresence_penalty,seed,response_format, …) are accepted but ignored (reported viax-yagami-ignored) — none of the CLI engines expose them. OpenAInmust be 1.thinkingand a yagami-extensioneffort("low"…"max") are passed through where the harness supports them (see the provider table) and reported as ignored elsewhere. OpenAIreasoning_effortmaps ontoeffort.- Cost is reported only by harnesses that price their own turns (Claude, OpenCode); Codex reports token usage without cost.
Development
bun run typecheck && bun run test # unit tests (no tokens spent)
bun run smoke # live end-to-end through your real Claude CLI (tiny token cost)
bun run live:providers # live check across every installed harness (tiny token cost each)
make build # build dist/ onlyCodex sessions preserve proposed plan documents and stream reasoning summaries as they arrive, and completed items only fill missing text. A failed resume is reported as an error; it never silently opens an empty conversation. Only one send can run at a time, including during startup. Input and permission handlers receive an abort signal when the server resolves their request, their turn stops, or the session closes; hosts should dismiss the corresponding prompt. Close and reopen a failed session with its last ID to retry.
