cli-openai-proxy
v0.1.0
Published
Run agentic coding CLIs (Claude Code, Codex, Paperclip adapters) behind an OpenAI-compatible API. Use your existing CLI subscription with any OpenAI client.
Maintainers
Readme
cli-openai-proxy
One OpenAI-compatible endpoint in front of the agentic coding CLIs you already have installed.
Claude Code and Codex — the two Paperclip
adapters registered here — run behind a single /v1/chat/completions endpoint,
and adding another that matches a supported output format and prompt-injection
strategy is one registry entry. Any OpenAI client —
an SDK, an IDE plugin, Collavre, your own service — can
drive them, from another machine if you want.
Bring your own key. The proxy owns no vendor account and issues no
credentials. Each CLI authenticates exactly as it does when you run it by hand,
with whatever credential you gave it — an API key, an OAuth login, a plan that
covers CLI usage. The proxy just spawns the CLI and translates its output. If a
CLI is not logged in, the completion fails with a 401 naming the engine to
re-authenticate; you can also supply the credential over HTTP through the
remote auth API.
Why
- CLI agents, not just chat models. These are agentic CLIs — they plan, call tools, and write files. This exposes that behind an interface every LLM client already speaks, instead of a bespoke integration per CLI.
- One endpoint, several engines. Switch engines by changing the
modelstring. Callers need no per-vendor SDK, key, or code path. - Remote by design. Run the proxy on the machine where the CLIs are installed and authenticated; call it from anywhere on your network. Remote auth provisioning means a client can recover from an expired login without anyone shelling into the host.
- BYOK, on your host. Credentials stay where you put them. Nothing is proxied through a third-party service.
Quick Start
# Install globally
npm install -g cli-openai-proxy
# Start it (needs at least one supported CLI installed and logged in)
cli-openai-proxy &
# See the supported model catalog (a static list — not a health check)
curl http://localhost:3456/v1/modelsOr clone and run:
git clone https://github.com/sh1nj1/cli-openai-proxy.git
cd cli-openai-proxy
npm install && npm run build
npm startSend a request. Pick the model that matches the CLI you actually have —
paperclip/codex_local for Codex, paperclip/claude_local for Claude Code:
curl -X POST http://localhost:3456/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "paperclip/codex_local", "messages": [{"role": "user", "content": "Hello!"}]}'How It Works
Your app / Collavre / IDE plugin (any OpenAI client)
|
| POST /v1/chat/completions { "model": "paperclip/codex_local", ... }
v
cli-openai-proxy <-- 127.0.0.1:3456 by default
|
| model id -> Paperclip adapter
v
Agent CLI as a subprocess claude | codex | ...
|
| your own credential on this host (BYOK)
v
Vendor backend --> CLI output --> OpenAI format (SSE or JSON) --> Your appEvery request runs through a Paperclip adapter, in a fresh temporary working directory that is removed afterwards: no CLI session and no working directory carries over between requests. That is a cwd, not a sandbox — the CLIs run with approvals bypassed, so a request that tells the agent to touch a path outside its cwd (your home directory, a repo on disk) still does, and the next request can see it. See docs/paperclip-adapters.md.
Engines and Models
Model ids follow one rule:
paperclip/<adapter>[/<cli-model>]| model value | Runs | Auth engine |
|---------------|------|-------------|
| paperclip/claude_local | Claude Code (claude), CLI's default model | claude |
| paperclip/claude_local/opus | Claude Code with --model opus | claude |
| paperclip/claude_local/claude-opus-4-6 | Claude Code with --model claude-opus-4-6 | claude |
| paperclip/codex_local | Codex (codex), CLI's default model | codex |
| paperclip/codex_local/gpt-5.4-mini | Codex with --model gpt-5.4-mini | codex |
| paperclip/codex_custom/anthropic/claude-sonnet-4.5 | Codex against a custom gateway (e.g. OpenRouter) | codex_custom |
The adapter part is matched against the registry; everything after it is the
model, passed to the CLI verbatim. The proxy keeps no model catalog of its own,
so a model works the moment your CLI supports it, with no proxy release — and an
id the CLI rejects surfaces as the CLI's own error rather than silently running
something else. Omit <cli-model> to let the CLI pick its default; omit model
from the request entirely and you get paperclip/claude_local.
Whatever the CLI takes for its own --model flag works here, because that is
literally where the value goes. For claude that is either an alias for the
latest of a family (fable, opus, sonnet, haiku) or a full id
(claude-opus-4-6); codex takes full ids only. Run claude --help for the
aliases your installed CLI knows — the proxy does not track them. Pick by intent, not by
correctness — an alias follows the CLI to each new model, a full id pins the one
you tested against.
An id outside this namespace returns 404 model_not_found. That includes every
id 1.x accepted — claude-opus-4, claude-max/…, anthropic/…,
claude-code-cli/…, and bare opus/sonnet/haiku — so a client configured
against those needs its model id updated before moving to 2.0; that break is
why this is a major version. Aliases are only meaningful as the
<cli-model> part, where the CLI resolves them: paperclip/claude_local/opus
runs, bare opus is a 404.
The model field on a response echoes the id you requested, unchanged, on both
the streaming and non-streaming paths — so a gateway that routes or validates on
it always sees an id this proxy accepts. The one exception is a missing or
falsy model (absent, "", null): the request falls back to the default
adapter and the response reports that default id, not your literal input.
This table — and GET /v1/models — is the set of adapters the proxy accepts,
not what this host can currently run: neither checks whether the underlying CLI
is installed or logged in, so a request can be accepted here and still fail at
execution. GET /v1/auth/{engine}/status (requires AUTH_ADMIN_KEYS; see
below) is an auth diagnostic, not an availability check: it never tests that the
CLI can run, and on the ordinary Claude path — logged in on the host, credential
in the keychain — it answers unknown, the same answer it gives when claude
is not installed at all. The reliable test is a request.
Adding an engine whose CLI matches an existing output format (Claude
stream-json or Codex JSONL) and an existing prompt-injection strategy
(task-context or prompt-template) is one registry entry in
src/adapter/paperclip-registry.ts plus its
published adapter package; a CLI that differs on either axis also needs the
matching parser/OutputMode or injection strategy added to PaperclipRunner
first.
Use with Collavre
Collavre models an AI agent as a user with an LLM vendor and a gateway URL, which is exactly the shape this proxy fits — no Collavre-side code, no API key.
- Run the proxy on a host that has the CLIs logged in, bound so Collavre can reach it (see Remote access).
- In the agent's settings, set vendor to
openaiand Gateway URL tohttp://<host>:3456/v1. - Set the model to the engine you want — e.g.
paperclip/codex_localorpaperclip/claude_local. - Leave the API key empty unless the proxy runs with
API_KEYS; then use one of those keys.
Collavre accepts arbitrary model ids on an OpenAI-compatible gateway, so each agent can point at a different CLI while sharing one proxy. Because the CLIs are agentic, a Collavre agent gets real tool use rather than plain completions. Full walkthrough: docs/paperclip-adapters.md.
Features
- OpenAI-compatible API — drop-in for any OpenAI client
- Any CLI model — the model part of the id is passed to the CLI verbatim, so the proxy keeps no model catalog to fall out of date
- Streaming — SSE deltas as the CLI produces output. Send
"stream_options": {"include_usage": true}to get the OpenAI terminal usage chunk (emptychoices) just before[DONE] - Cache-aware token counts —
prompt_tokenscovers the whole prompt, cache reads included, with the cached share broken out asusage.prompt_tokens_details.cached_tokens - Image input — OpenAI
image_urlparts (base64 data URLs) are materialized to temp files and handed to the CLI as inline links, so the agent can see them; works across all adapters - OpenAI-shaped errors — usage limits become
429 insufficient_quota, an unauthenticated CLI becomes401 engine_unauthenticated.Retry-Afterrides along only when the adapter reported a reset time; otherwise the429has no retry delay to advertise. Those statuses and headers apply to non-streaming requests; a streaming request has already flushed200, so the same classified error arrives in-band as an SSE{"error": {...}}object carrying the sametype/code - Remote CLI auth provisioning — log a CLI in over HTTP
- Usage tracking — token counts, latency, and request history. The stored
model label is bucketed to
opus/sonnet/haikufor cost estimation, so it does not identify which engine served a request - API key auth — optional Bearer tokens for shared deployments
- Per-user Linux workers — first request creates a non-login OS account and socket-activated worker; CLI processes, HOME, credentials, and usage stay per user
- Docker deployment — one image for production and CI, giving macOS and Windows hosts the same per-user Linux isolation via Docker Desktop's Linux VM; see docs/docker.md
- Stateless execution — fresh isolated workspace per request
- Auto-start — macOS LaunchAgent for an always-on service
Configuration
| Env var | Default | Meaning |
|---------|---------|---------|
| PORT | 3456 | Listen port (also accepted as the first CLI argument) |
| HOST | 127.0.0.1 | Bind address — set 0.0.0.0 for remote access |
| API_KEYS | (unset) | Comma-separated Bearer tokens for callers. Unset = open access |
| AUTH_ADMIN_KEYS | (unset) | Separate key set gating /v1/auth/*. Unset = those routes are disabled |
| AUTH_TRUST_COMPLETION_CALLERS | (unset) | Declares completion callers trusted with a provisioned Claude credential |
| PROVISION_SYNC | (unset) | Enables the /v1/provision/* agent-provisioning routes; see docs/provisioning.md |
| USER_WORKER_MODE | (unset) | Route user-scoped APIs to isolated OS-user workers |
| USER_API_KEYS | (unset) | JSON mapping of opaque caller keys to stable tenantId + userId |
| USER_IDENTITY_HMAC_SECRET | (unset) | Shared secret for trusted upstream identity headers |
| PROXY_REQUIRE_IDENTITY_V2 | (unset) | Require signed workspace-aware v2 identities; mapped user keys remain supported |
| PROVISION_WORKSPACE_ROOT | ~/workspaces | Parent directory for named agent workspaces in worker mode |
| PROVISION_MAX_WORKSPACES_PER_USER | 32 | Bound named workspaces per user/worker |
| USER_WORKER_PROVISIONER_ENDPOINT | platform default | Provisioner Unix socket or Windows Named Pipe |
| TIMEOUT | 0 (none) | Per-request ceiling in ms |
| DEBUG | (unset) | Verbose logging |
Remote access
HOST=0.0.0.0 makes the proxy reachable from other machines. Set API_KEYS
whenever you do: the CLIs run with approvals bypassed, so an unauthenticated
caller effectively has code execution on the host.
API_KEYS authenticates callers but does not encrypt anything. The proxy speaks
plain HTTP, so the Bearer key and every prompt cross the network in the clear —
anyone who can observe the traffic can replay the key. Beyond loopback, put it
behind TLS (a reverse proxy terminating HTTPS) or an encrypted tunnel such as
Tailscale or WireGuard; the plain http:// examples below assume a trusted link.
HOST=0.0.0.0 API_KEYS=sk-team-abc123 cli-openai-proxy
PROXY_HOST=proxy.internal.example
curl "http://$PROXY_HOST:3456/v1/chat/completions" \
-H "Authorization: Bearer sk-team-abc123" \
-H "Content-Type: application/json" \
-d '{"model": "paperclip/claude_local", "messages": [{"role": "user", "content": "Hello!"}]}'Remote CLI Auth Provisioning
Log the underlying claude / codex CLIs in over HTTP instead of shelling into
the host — so a remote UI can recover from an auth failure on its own. Off unless
AUTH_ADMIN_KEYS is set (its own key set, separate from API_KEYS):
Start the server:
API_KEYS=sk-team-abc123 \
AUTH_ADMIN_KEYS=sk-admin-xyz789 \
AUTH_TRUST_COMPLETION_CALLERS=1 \
cli-openai-proxyThen, in another terminal:
# 1. open a session — codex returns an API-key prompt, claude an OAuth URL
SESSION_ID=$(curl -sX POST -H "Authorization: Bearer sk-admin-xyz789" \
http://localhost:3456/v1/auth/codex/sessions | jq -r .sessionId)
# 2. submit the API key (claude: the code from the OAuth URL) to finish
curl -X POST -H "Authorization: Bearer sk-admin-xyz789" \
-H "Content-Type: application/json" \
-d '{"value": "sk-..."}' \
http://localhost:3456/v1/auth/codex/sessions/$SESSION_IDAUTH_TRUST_COMPLETION_CALLERS=1 is required for the claude and codex_custom
engines. It declares that every completion-key holder is trusted with the
provisioned credential — those CLIs read it from their environment, so the proxy
has to hand it to a child any completion caller can inspect. Without it, the
session request returns 403 caller_trust_not_declared. codex does not need it:
codex login persists its own credential and nothing is injected.
A completion whose CLI is unauthenticated answers 401 with
code: "engine_unauthenticated" and the engine to re-authenticate, so a client
can trigger the right flow automatically.
Custom gateways (OpenRouter and other OpenAI-compatible endpoints)
paperclip/codex_custom runs the same codex CLI against an endpoint you
provision, on that endpoint's own API key. Both values are submitted together —
a key alone does not say where to spend it:
SESSION_ID=$(curl -sX POST -H "Authorization: Bearer sk-admin-xyz789" \
http://localhost:3456/v1/auth/codex_custom/sessions | jq -r .sessionId)
curl -X POST -H "Authorization: Bearer sk-admin-xyz789" \
-H "Content-Type: application/json" \
-d '{"api_key": "sk-or-v1-...", "base_url": "https://openrouter.ai/api/v1"}' \
http://localhost:3456/v1/auth/codex_custom/sessions/$SESSION_IDThen name a model the gateway serves:
{"model": "paperclip/codex_custom/anthropic/claude-sonnet-4.5"}. There is no
default worth inheriting here — the CLI's own default model is an OpenAI id the
gateway probably does not carry — so pass the <cli-model> part.
Requirements and behaviour:
- The gateway must serve OpenAI's Responses API (
POST <base_url>/responses). Codex ≥ 0.145 refuses to load a config asking for Chat Completions, so a Chat-only gateway cannot be used. OpenRouter serves both. base_urlmust behttps, except on loopback: the key travels to it as a bearer token on every request.- The key is held in proxy memory only and reaches the CLI as
CODEX_CUSTOM_API_KEY. It is never written to disk — the generatedconfig.tomlreferences the env var rather than the value. - This engine is independent of
paperclip/codex_local, which keeps running on yourcodex login. They use separateCODEX_HOMEdirectories, so provisioning a gateway never disturbs a ChatGPT subscription login.
See docs/cli-auth-provisioning.md.
API Endpoints
| Endpoint | Method | Description |
|----------|--------|-------------|
| /health | GET | Health check + usage summary |
| /v1/models | GET | Supported model catalog — static, does not check whether a CLI is installed or authenticated |
| /v1/chat/completions | POST | Chat completions (streaming & non-streaming) |
| /v1/usage | GET | Usage stats |
| /v1/usage/recent | GET | Recent request log |
| /v1/auth/engines | GET | Auth flow per engine (needs AUTH_ADMIN_KEYS) |
| /auth | GET | Browser UI for CLI authentication (available when AUTH_ADMIN_KEYS is set) |
| /v1/auth/{engine}/status | GET | Whether that CLI is authenticated |
| /v1/auth/{engine}/sessions | POST | Start a login flow |
| /v1/auth/{engine}/sessions/{id} | GET / POST / DELETE | Poll / submit / abandon |
| /v1/auth/{engine}/credential | DELETE | Forget a provisioned credential |
| /v1/provision | GET | Agent provisioning status (needs PROVISION_SYNC=1 + AUTH_ADMIN_KEYS) |
| /v1/provision/manifest | POST | Register a manifest URL and apply it, without a login |
| /v1/provision/sync | POST | Re-fetch the provisioning manifest and apply it |
| /v1/provision/items/{type}/{name}/approve | POST | Approve a first-seen provisioned item |
| /v1/provision/items/{type}/{name} | DELETE | Uninstall an item and revoke its approval |
Integration Examples
Python (OpenAI SDK)
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:3456/v1",
api_key="not-needed", # or one of API_KEYS
)
response = client.chat.completions.create(
model="paperclip/codex_local",
messages=[{"role": "user", "content": "Hello!"}],
)Continue.dev / Cursor
{
"models": [{
"title": "Claude Code CLI",
"provider": "openai",
"model": "paperclip/claude_local/opus",
"apiBase": "http://localhost:3456/v1",
"apiKey": "not-needed"
}]
}OpenClaw
Configure an OpenAI-compatible provider pointing at localhost:3456, with the
model set to a paperclip/… id. The built-in claude-max provider preset sends
claude-max/claude-* ids, which this proxy does not accept.
cURL (streaming)
curl -N -X POST http://localhost:3456/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "paperclip/claude_local", "messages": [{"role": "user", "content": "Hello!"}], "stream": true}'Usage Tracking
Token counts and latency per request, persisted to ~/.cli-openai-proxy/usage.json:
curl http://localhost:3456/v1/usage
{
"totalRequests": 847,
"totalInputTokens": 12500000,
"totalOutputTokens": 3200000,
"estimatedApiCostSavedUsd": 427.50,
"avgResponseMs": 2340,
"byModel": {
"opus": { "requests": 523, "estimatedCostUsd": 389.20 },
"sonnet": { "requests": 324, "estimatedCostUsd": 38.30 }
},
"maxSubscriptionCostUsd": 200
}
# Recent requests
curl http://localhost:3456/v1/usage/recent?limit=10estimatedApiCostSavedUsd and maxSubscriptionCostUsd are priced against
Anthropic's published API rates and a fixed $200 figure — leftovers from the
project's Claude-only origins. Treat them as a rough reference for Claude
traffic only; they say nothing about other engines.
Counts follow what the CLI reports it ran, not the requested id — paperclip/claude_local
names an adapter, so the model is only known after the run. Subagent tokens are
included, and a turn spanning several models is priced per model while staying one
request under the model that produced most of its output.
Prerequisites
- Node.js >= 22.13.0
- At least one supported agent CLI, installed and authenticated with your own
credential:
npm install -g @anthropic-ai/claude-code # then: claude (log in) npm install -g @openai/codex # then: codex login
The proxy starts even if a CLI is missing — a request targeting that engine fails at request time with a clear error, so a codex-only or claude-only host is fine.
Auto-Start on macOS
cat > ~/Library/LaunchAgents/com.cli-openai-proxy.plist << 'EOF'
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.cli-openai-proxy</string>
<key>ProgramArguments</key>
<array>
<string>/usr/local/bin/node</string>
<string>/path/to/cli-openai-proxy/dist/server/standalone.js</string>
</array>
<key>RunAtLoad</key>
<true/>
<key>KeepAlive</key>
<true/>
</dict>
</plist>
EOF
launchctl load ~/Library/LaunchAgents/com.cli-openai-proxy.plistSee docs/macos-setup.md.
Auto-Start on Ubuntu
Run the installer as the same user that authenticated the agent CLI(s). Do not
run it with sudo.
git clone https://github.com/sh1nj1/cli-openai-proxy.git
cd cli-openai-proxy
./scripts/install-linux-single-user.shThe installer reuses a trusted Node.js 22.13.0 or newer, or downloads and
checksum-verifies the latest supported Node.js 22 runtime under /opt when none is
available. It installs dependencies, builds the project, and enables the
com.cli-openai-proxy.service systemd user service. It also installs the
engine CLIs (INSTALL_CLIS, default @anthropic-ai/claude-code @openai/codex)
into a managed prefix under ${XDG_DATA_HOME:-$HOME/.local/share} on the
service PATH; CLIs already on your PATH take precedence, and INSTALL_CLIS=""
skips the step. It also enables
systemd linger so the proxy starts at boot before login. On a minimal Ubuntu
installation it also installs build-essential and Python 3, which are needed
to compile node-pty. sudo is used only for those OS prerequisites and the
linger setting; the proxy itself runs as the current user.
The default listener is 127.0.0.1:3456. Edit
${XDG_CONFIG_HOME:-$HOME/.config}/cli-openai-proxy.env to configure
PORT, HOST, API_KEYS, or remote CLI authentication, then restart the
service. The installer also prints the effective path as Config:. If HOST
is changed to a non-loopback address, set API_KEYS before exposing the port.
The service PATH includes common user CLI locations and trusted absolute
entries from PATH when the installer runs. Directories writable by users other
than the service user are skipped. Rerun ./scripts/install-linux-single-user.sh after adding a new
custom CLI installation directory to PATH.
systemctl --user restart com.cli-openai-proxy.service
systemctl --user status com.cli-openai-proxy.service
journalctl --user -u com.cli-openai-proxy.service -f
curl http://127.0.0.1:3456/healthAfter pulling new code from main, rerun ./scripts/install-linux-single-user.sh. Existing environment
configuration is preserved while dependencies, the build, and the service are
updated.
Multi-user Linux service
For a shared gateway where each authenticated caller's CLI must run as a different Linux account, use the root provisioner + per-user worker installer:
sudo ./scripts/install-linux-user-workers.shThis is a one-shot install: it prepares a root-trusted Node.js runtime, runs
npm ci and the build as a dedicated build-only non-login account, validates
the production dependencies, and promotes only that frozen result into a
root-owned release under /opt. It then installs and starts the provisioner,
gateway, and worker systemd units and verifies the gateway health endpoint.
On first install it writes separate random USER_API_KEYS and
AUTH_ADMIN_KEYS values to /etc/cli-openai-proxy/gateway.env, creates a
default/default user mapping, and prints the new keys once. Override that
mapping with INSTALL_TENANT_ID and INSTALL_USER_ID. Reinstalls preserve the
configuration and keys; use INSTALL_ROTATE_KEYS=1 only when intentional.
When converting from the Single installer, run via sudo from that service
user (or set INSTALL_SINGLE_USER) so the conflicting user service is stopped
and disabled before port 3456 moves to Multi mode. Per-user CLI credentials are
not migrated: repeat Codex device login and agent provisioning for each mapped
user after installation.
The first request creates a locked, non-login Linux account and starts its
systemd socket. The public gateway remains a dedicated low-privilege service
account, and the root provisioner executes only the root-owned runtime copied
under /opt. See Per-user Linux workers for identity
signing, systemd configuration, auth provisioning, and the multi-OS adapter
boundary.
Architecture
src/
├── adapter/ # OpenAI <-> CLI conversion, adapter registry, output parsers
├── auth/ # Remote CLI login flows, pty driver, in-memory token store
├── cli/ # CLI presence checks
├── isolation/ # User identity, IPC proxy, OS provisioner adapters
├── server/ # Express server, routes, proxy auth
├── usage/ # Token tracking and analytics
└── types/ # TypeScript type definitionsSecurity
- CLIs are spawned via
spawn(), never a shell (no injection) - Every run gets a fresh temporary working directory
- CLIs run with approvals bypassed — anyone who can call
/v1/chat/completionscan run code on the host. SetAPI_KEYSon any non-loopback bind - Proxy access keys and user-identity secrets are captured before startup preflight and removed from every CLI subprocess environment
- Per-user mode never trusts the OpenAI
userfield, and never falls back to a shared account when identity, provisioning, or worker connection fails - Codex provisioning forwards an API key to
codex loginover stdin; the CLI persists it in~/.codex - Claude provisioning captures the
setup-tokenOAuth credential, keeps it in proxy memory only, and injects it only when the completion-caller trust boundary is explicitly accepted
Important Disclaimer
Completions run the official vendor CLIs as subprocesses; the proxy does not reverse-engineer private APIs or bypass authentication. The optional remote-auth API does handle credentials — review the credential exposure and trust model before enabling it.
Each vendor sets its own terms for how its CLI may be used, including from automation. Review the terms for the CLIs you enable (Anthropic, OpenAI) before deploying. Policies on third-party tooling may change. Use at your own discretion and risk.
Contributing
PRs welcome. Please include tests.
Credits
Originally created by Atal Ashutosh as atalovesyou/claude-max-api-proxy. This repository continues that work with multi-CLI adapters (Claude Code, Codex, Paperclip), streaming, usage tracking, and remote auth provisioning.
License
MIT — see LICENSE. Copyright (c) 2026 Atal Ashutosh.
