npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

aio-proxy

v0.39.0

Published

All-in-one LLM API proxy

Downloads

6,467

Readme

Version Downloads License

English | 简体中文 | Documentation

AIO Proxy is a local model gateway. Keep the SDKs and coding agents you already use, plug in API keys or the subscriptions you already pay for, and get cross-protocol conversion, routing with automatic failover, and full request traces — all from one binary on one port.

For example: Claude Code speaks Anthropic Messages, your ChatGPT subscription speaks OpenAI Responses, and a backup API key speaks Gemini. Point Claude Code at AIO Proxy and it just works — tool calls, reasoning, and streaming are converted on the fly, and when one upstream fails the request falls through to the next.

flowchart LR
  subgraph Clients["Your clients, unchanged"]
    Agents["Coding agents<br/>Codex · Claude Code · OpenCode · Pi · Grok Build"]
    SDKs["SDKs and apps<br/>OpenAI · Anthropic · Gemini"]
  end

  Proxy["AIO Proxy<br/>Protocol conversion · Routing & failover<br/>Usage, cost & traces"]

  subgraph Upstreams["Any upstream"]
    Keys["API keys<br/>OpenAI · Anthropic · Gemini · any compatible API"]
    Subs["Subscriptions via OAuth<br/>ChatGPT · Claude · Copilot · Cursor · Grok …"]
    SDKProviders["AI SDK provider packages"]
  end

  Agents --> Proxy
  SDKs --> Proxy
  Proxy --> Keys
  Proxy --> Subs
  Proxy --> SDKProviders

Why AIO Proxy

  • Every protocol, one endpoint: OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, and Gemini clients all share one port. Matching protocols pass through raw; everything else is converted, including tools, reasoning, and streaming. Embeddings, images, audio, and token counting are covered too.
  • Bring your subscriptions: Log in with OAuth to ChatGPT, Claude, GitHub Copilot, Google Antigravity, Cursor, xAI Grok, Kimi Code, OpenRouter, and more, and use them as standard API endpoints.
  • The whole AI SDK ecosystem: Any Vercel AI SDK provider package — official or community — loads as a Provider with kind: "ai-sdk" and gets the same conversion, routing, and billing as everything else. If the AI SDK supports a vendor, so does AIO Proxy.
  • Extend it with plugins: Every built-in subscription above is a plugin built on the public @aio-proxy/plugin-sdk. Write your own for an unsupported service or an internal gateway — OAuth login, model catalog, metadata and pricing included — and install it with aio-proxy plugin add.
  • Routing that survives outages: Provider priority tiers decide who is tried first, Provider weight splits traffic within a tier, session affinity keeps prompt caches warm, and a failed upstream falls through to the next candidate. Priority, weight, price, and context limits can all be overridden per model, with aliases to unify names.
  • Coding agents in one command: aiop agent configure wires up Codex, Claude Code, Grok Build, OpenCode, Pi, and oh-my-pi, with device approval instead of pasted keys where the Agent supports it; anything else only needs a base URL.
  • Requests you can see: The built-in Dashboard records every request and every Provider attempt with status, latency, tokens, and cost, with full traces exportable to OpenTelemetry.
  • Local-first, configured your way: Binds to 127.0.0.1 by default; add caller API keys and a Dashboard password when you expose it. Each Provider can declare multiple protocol endpoints, custom headers, and its own HTTP(S)/SOCKS5 proxy with fallback. Configure in the Dashboard or in a schema-checked JSONC file with {{env.NAME}} secrets and hot reload.

Reuse a sign-in already on this machine

ChatGPT and GitHub Copilot Providers can reuse the sign-in their tools already keep on this machine. In the Dashboard, choose Use the Codex sign-in on this machine or Use the GitHub Copilot sign-in on this machine beside the browser authorize button. The option appears only when a sign-in is detected on the machine running the server; containers and headless servers without that sign-in show no option. Codex with cli_auth_credentials_store = "keyring" also shows no option because it has no sign-in file.

From the CLI, run aio-proxy provider login --local-sign-in, or choose the local sign-in in the interactive aio-proxy provider login prompt. Credentials are read only after you choose this option. aio-proxy writes rotated tokens back to keep Codex signed in when it refreshes; the Copilot token is read once at link time and does not rotate.

Linked Providers show a Linked to {source} on this machine badge. Removing the Provider never signs the tool out. To recover:

  • If the Provider is disabled while the tool still works, use the local sign-in again on that Provider.
  • If both are signed out, sign in again in the tool, then use the local sign-in again.
  • If the tool is on a different account, switch it back or add a new Provider.

Install

Homebrew

brew install aio-proxy/tap/aio-proxy

Bun

bun add -g aio-proxy

curl

curl -fsSL https://aioproxy.dev/install.sh | sh

Quick start

aio-proxy run --open

Every install method (Homebrew, npm/bun, and the curl installer) also provides the short command aiop.

  • API: http://127.0.0.1:9317
  • Dashboard: http://127.0.0.1:9317/dashboard

The first run creates ~/.aio-proxy/config.jsonc. The initial configuration has no Providers; add one through the Dashboard or edit the configuration file directly:

aio-proxy config path
aio-proxy config edit

Configuration

The following example routes gpt-5 to the OpenAI Responses API:

{
  "$schema": "https://unpkg.com/@aio-proxy/types/config.schema.json",
  "providers": {
    "openai": {
      "kind": "api",
      "protocol": "openai-response",
      "baseURL": "https://api.openai.com/v1",
      "apiKey": "{{env.OPENAI_API_KEY}}",
      "models": ["gpt-5"],
    },
  },
}

Validate or reload the configuration:

aio-proxy config validate
aio-proxy reload

Editors that support $schema can provide completion and validation. Use {{env.NAME}} to read environment variables.

Sync models with upstream

API and AI SDK Providers can opt into syncModels: true instead of maintaining models. Sync is off by default; enabling it together with a non-empty models list fails validation. excludedModels is only valid with syncModels: true and hides exact model IDs (no glob matching). Aliases can still target hidden models. A manual models list works as before.

{
  "providers": {
    "relay": {
      "kind": "api",
      "protocol": "openai-compatible",
      "baseURL": "https://relay.example.com",
      "apiKey": "{{env.RELAY_API_KEY}}",
      "syncModels": true,
      "excludedModels": ["gpt-3.5-turbo"],
    },
  },
}

Discovery runs immediately on startup and after config changes, then every hour. Failed refreshes retry after 5 minutes once the list is stale; an outage or empty response keeps the last good list. Before the first successful discovery, only aliases route. Changing the API primary endpoint's baseURL, protocol, or endpoint form, or the AI SDK packageName or options.baseURL, discards the old list. Only the primary API endpoint is queried, and discovery never rewrites the config file.

AI SDK sync requires the package instance's listModels method or options.baseURL serving an OpenAI-compatible /models endpoint; otherwise the Provider reports CATALOG_UNSUPPORTED. In the Dashboard models section, choose Manual / Sync with upstream, hide individual models, check the last refreshed time, or use the refresh button.

Multi-protocol endpoints

Some upstreams natively serve more than one protocol. Declare the extra endpoints with endpoints; a request whose inbound protocol matches any declared endpoint is forwarded verbatim (raw passthrough) instead of being converted:

{
  "providers": {
    // First-party channel: per-protocol endpoints. Keep the legacy pair as the
    // primary endpoint and append the extra protocols.
    "moonshot": {
      "kind": "api",
      "protocol": "openai-compatible",
      "baseURL": "https://api.moonshot.cn/v1",
      "apiKey": "{{env.MOONSHOT_API_KEY}}",
      "models": ["kimi-k2"],
      "endpoints": [{ "protocol": "anthropic", "baseURL": "https://api.moonshot.cn/anthropic/v1", "auth": "bearer" }],
    },
    // Aggregator gateway: one AI SDK-style base URL shared by several protocols.
    "gateway": {
      "kind": "api",
      "apiKey": "{{env.GATEWAY_KEY}}",
      "models": ["gpt-5"],
      "endpoints": { "baseURL": "https://gw.example.com/v1", "protocol": ["openai-response", "anthropic"] },
    },
  },
}

Rules:

  • An endpoints entry's baseURL is exactly what you would pass to the matching AI SDK package: OpenAI-style and Anthropic endpoints include the /v1 segment, Gemini endpoints include /v1beta (so Gemini cannot share a /v1 base URL — give it its own array entry).
  • Vendor docs often quote the Anthropic base for ANTHROPIC_BASE_URL (for example https://api.z.ai/api/anthropic); append /v1 when copying it here.
  • auth is only supported on anthropic endpoints (declaring it on an endpoint of any other protocol fails validation): bearer sends Authorization: Bearer and requires the provider to declare apiKey, the default x-api-key keeps today's header.
  • The top-level protocol/baseURL pair stays the primary endpoint and keeps its historical passthrough behavior — on passthrough its base URL's path is discarded and only the origin is used, joined with the inbound request path, so services with a required path prefix should use endpoints; cross-protocol conversion always targets the primary endpoint. Without a top-level pair, the primary endpoint is the first endpoints entry (in the shared form, the first protocol in its protocol list).
  • The Dashboard supports shared and independent endpoint addresses. New API Providers use endpoints, including when only one protocol is selected, so the full upstream path is preserved. Existing legacy single-protocol Providers retain their top-level protocol/baseURL behavior.

Command Code

See the Command Code integration guide for API-key setup, subscription eligibility, and separate OpenAI-compatible and Anthropic Provider examples that preserve the /provider/v1 path.

Model metadata and pricing

Configure client-facing metadata once per exposed model under router.models.<slug>.metadata, keyed by the exact slug clients request rather than by an upstream model id. The slug must already be exposed by a Provider's models, synced catalog, or alias configuration: a router.models entry only customizes an existing route and never creates one. The removed providers.<id>.metadata field is silently ignored.

Metadata is resolved per field in this order: the selected Provider's router override (for cost or limit) > slug metadata (including extend) > plugin-reported upstream metadata > models.dev fallback > protocol default. A Provider override replaces the slug's entire cost or limit object rather than deep-merging it; other metadata is shared by every Provider serving that slug. Aliases only auto-discover catalog fallback by their public slug. Unknown metadata fields are preserved and warned about rather than rejected, while invalid values (for example a negative price or a non-positive context limit) fail validation with a clear error.

{
  "$schema": "https://unpkg.com/@aio-proxy/types/config.schema.json",
  "router": {
    // When several Providers expose the same public model, reconcile its context window:
    // "min" (default, safe) reports the smallest; "max" reports the largest.
    "modelContextAggregation": "min",
    "models": {
      // Keyed by the exposed slug clients request.
      "gpt-5": {
        "metadata": {
          "name": "GPT-5", // client-facing display name
          "description": "Frontier model",
          "limit": {
            "context": 400000,
            "input": 272000,
            "output": 128000,
          },
          "capabilities": {
            "reasoning": true,
            "toolCall": true,
            "attachment": true,
            // This can grant Images support to this slug without affecting other slugs.
            "modalities": { "input": ["text", "image"], "output": ["text", "image"] },
          },
          "cost": {
            // Per-token prices are USD per 1,000,000 tokens.
            "input": 1.25,
            "output": 10,
            "cacheRead": 0.125,
            // Audio token prices, USD per 1,000,000 tokens (OpenAI-compatible upstreams only).
            "inputAudio": 2.5,
            "outputAudio": 20,
            // Per-event fees are USD per event.
            "image": 0.01,
            "webSearch": 0.01,
            "request": 0,
            // Long-context surcharge: the highest crossed tier applies to the whole request.
            "tiers": [{ "tier": { "type": "context", "size": 200000 }, "input": 2.5, "output": 15 }],
          },
        },
        "providers": {
          // Provider-specific values replace metadata.cost and metadata.limit wholesale.
          "openai": {
            "cost": { "input": 1, "output": 8 },
            "limit": { "context": 300000, "input": 200000, "output": 100000 },
          },
        },
      },
    },
  },
  "providers": {
    "openai": {
      "kind": "api",
      "protocol": "openai-response",
      "baseURL": "https://api.openai.com/v1",
      "apiKey": "{{env.OPENAI_API_KEY}}",
      "models": ["gpt-5"],
    },
  },
}

limit.context is the maximum total context, limit.input is the maximum input tokens, and limit.output is the maximum output tokens. Configured input and output cannot exceed configured context. For Codex, these distinct limits project to context_window = input ?? context and max_context_window = context ?? input; output is never used as a Codex context window.

When a request is billed, the Provider that actually served it supplies the price: its router cost override wins over slug metadata and the models.dev catalog, and the recorded usage row notes whether the price came from config, models-dev, or a built-in default (priceSource).

Per-event fees and audio-token costs are metered from the actual response: generated images and web-search invocations are counted from the served output, and audio tokens are read from the upstream usage (available on OpenAI-compatible Chat Completions upstreams). A fee applies only when the corresponding events occur.

Inheriting a catalog entry with extend

When an exposed slug doesn't line up with a models.dev slug (for example, an aliased or renamed model), point extend at the catalog slug to inherit as a base layer:

{
  "router": {
    "models": {
      // This exposed slug must also be listed by a Provider or produced by an alias.
      "my-frontier-alias": {
        "metadata": {
          "extend": "openai/gpt-5.5", // inherit this catalog entry as the base
          "name": "My Frontier Model", // override the inherited name
          "cost": { "input": 2 }, // override input price; inherited output/tiers remain
        },
        "providers": {},
      },
    },
  },
}
  • extend names a models.dev slug (provider/model) whose catalog entry supplies the base metadata (name, limit, capabilities, cost).
  • Your explicit fields override the inherited ones. Merging is deep for objects (e.g. cost.input above overrides only that field while cost.output is inherited), and arrays (such as capabilities.reasoningOptions, modalities, and cost tiers) replace the inherited array wholesale rather than merging by index.
  • Only the extend target is used as the base — the model's own upstream id is not auto-matched against the catalog. That is the whole point of extend: the name doesn't line up.
  • Inherited cost is treated as a config price: you opted in through extend, so billing tags it priceSource: "config", just like a cost you wrote out in full.
  • If the target slug isn't found in the catalog, your explicit fields are kept (the extend key is dropped) and a warning is logged; startup is never blocked.

Routing rules

Each key in the providers object is a stable Provider ID. Provider priority is an integer failover tier (0..10000, default 0); higher values are tried first. Provider weight distributes traffic within one priority tier: it is a finite authored number, default 1, then Math.round and clamped to 0..10000. Existing configurations keep their old weight values, but that field no longer defines a global fixed order.

providers:
  provider-a:
    priority: 0
    weight: 1000
router:
  models:
    model-m:
      providers:
        provider-a: { priority: 30, weight: 6000 }
        provider-b: { priority: 30, weight: 4000 }
        provider-c: { priority: 20 }

router.models keys are exact client-requested model IDs. They never create a candidate, choose an upstream target, or use glob matching. A missing Provider entry or field inherits the Provider default. A positive model weight can re-enable a Provider whose default weight is zero.

A request is handled as follows:

  1. Try the complete request model string as an exact Provider-qualified route first. If it matches, select that Provider directly and bypass Provider priority and Provider weight, including effective weight zero. enabled: false still blocks the Provider because disabled Providers are not in the route map.
  2. Otherwise try the same complete string as an exact normal client model ID, including strings containing /.
  3. Merge Provider defaults with the exact model's sparse providers overrides. Discard normal candidates with enabled: false or effective weight zero.
  4. Order remaining candidates by descending Provider priority, then by Provider weight within the same priority tier. With the opt-in router.selection: "quota-reset" (also a switch on the Dashboard Routing page), subscriptions with a known quota reset go first within their tier, the one whose longest covering window resets soonest leading; Providers without quota data follow in their weighted order. Configuration order is a deterministic tie-breaker for catalog representation and diagnostics, not the request order for positive-weight candidates in the same tier.
  5. Stable (non-generated) logical sessions use a deterministic weighted draw so token-count and generation share the same pre-attempt order when the routing snapshot is unchanged. Generated sessions use independent random draws.
  6. Response owner, then session affinity, may move an eligible normal candidate to the front. They never resurrect a disabled or zero-weight Provider. Session affinity still overrides priority so a session can stick to a previously successful Provider (for example, prompt-cache continuity).
  7. Remove candidates that cannot serve right now, even when response owner or session affinity put them first: a Provider cooling down after a 429 with Retry-After, and a subscription Provider whose cached quota snapshot shows the window covering this model exhausted with a known reset (Kimi Code, Muse Code, ChatGPT, and Cursor report which models each window covers). Quota that is unknown, failed to read, or older than 10 minutes never removes a candidate. If every candidate is removed, the client gets a 429 whose Retry-After is the earliest reset. The request trace lists removed candidates in aio_proxy.route.skipped_candidates.
  8. Use raw passthrough for a same-protocol api Provider; use AI SDK conversion for other supported combinations.
  9. Try the next candidate after a Provider failure; return the final failure if every candidate fails.

On the example policy, provider-a is first about 60% of the time and provider-b about 40% at priority 30. If the selected Provider fails, the other priority-30 Provider is tried before provider-c.

If every normal candidate is disabled or has effective weight zero, the model is omitted from GET /v1/models and normal requests use the existing model-unavailable/not-found behavior. An enabled Provider remains reachable through an exact Provider-qualified request.

GET /v1/models is deterministic even though request selection is weighted. The public representative Provider is chosen from enabled, positive-effective-weight candidates by highest Provider priority, then highest Provider weight, then original configuration order.

Preserving previous weight-as-order behavior

Previously, Provider weight was a global fixed order: unique weights were tried high-to-low, and equal or omitted weights preserved configuration order. Under this contract both Providers default to Provider priority 0, and Provider weight is a same-tier traffic share. Existing files are not rewritten.

| Old configuration | New configuration to preserve intent | | --------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- | | Unique old weights used as fixed order | Copy old weight to priority; set new weight: 1. | | Equal old weights whose config order mattered | Assign explicit descending priorities in the old config order; set weight: 1. | | Omitted old weight | It previously remained eligible at priority zero; set a positive new weight, normally 1. | | Fractional old weight | Assign priority from the old descending order; the new weight is rounded with Math.round if retained as a traffic ratio. | | Negative or greater-than-10000 old weight | Assign explicit in-range priorities that preserve the old total order; do not copy values that would collapse under clamp. | | Old weight: 0 | It previously remained an eligible fallback; set new weight: 1 and the intended priority. | | enabled: false | No change; it remains the hard disable. |

API

| Protocol or purpose | Method and path | | --------------------------- | --------------------------------------------------- | | Health check | GET /health | | Model list | GET /v1/models | | OpenAI Chat Completions | POST /v1/chat/completions | | OpenAI Responses | POST /v1/responses | | OpenAI Completions | POST /v1/completions | | OpenAI Responses compact | POST /v1/responses/compact | | Anthropic Messages | POST /v1/messages | | Anthropic Token Counting | POST /v1/messages/count_tokens | | Gemini | POST /v1beta/models/{model}:generateContent | | Gemini streaming | POST /v1beta/models/{model}:streamGenerateContent | | Gemini Token Counting | POST /v1beta/models/{model}:countTokens | | Gemini Interactions | POST /v1beta/interactions | | OpenAI Embeddings | POST /v1/embeddings | | Gemini embed | POST /v1beta/models/{model}:embedContent | | Gemini batch embed | POST /v1beta/models/{model}:batchEmbedContents | | OpenAI Images generations | POST /v1/images/generations | | OpenAI Images edits | POST /v1/images/edits | | OpenAI Audio speech | POST /v1/audio/speech | | OpenAI Audio transcriptions | POST /v1/audio/transcriptions | | OpenAI Audio translations | POST /v1/audio/translations | | Codex Live create | POST /v1/live | | Codex Realtime create | POST /v1/realtime | | Codex Realtime calls | POST /v1/realtime/calls | | Codex Live sideband | GET /v1/live/:call_id | | Codex Realtime sideband | GET /v1/realtime/calls/:call_id | | Codex Realtime WS | GET /v1/realtime | | Codex Realtime hangup | POST /v1/realtime/calls/:call_id/hangup | | OpenAI Videos create | POST /v1/videos | | OpenAI Videos retrieve | GET /v1/videos/:video_id | | OpenAI Videos content | GET /v1/videos/:video_id/content | | OpenAI Videos delete | DELETE /v1/videos/:video_id | | OpenAI Videos remix | POST /v1/videos/:video_id/remix | | OpenAI Videos edits | POST /v1/videos/edits | | OpenAI Videos extensions | POST /v1/videos/extensions | | TypeSafe System One | POST /v1/systemone |

Images notes:

  • Raw Images needs an openai-image endpoint (or primary protocol).
  • Missing/null/empty/whitespace JSON model and multipart missing/empty/whitespace model look up gpt-image-2.5-sunburst. Explicit model IDs, including gpt-image-2 and gpt-image-2.5-flare, keep their existing routing. Multipart literal form null is the model id "null" and is not defaulted. Raw injects the resolved candidate id.
  • Inbound model is the OpenAI id plus the existing providerId/ qualifier.
  • Convert does not stream and does not fetch image_url.
  • DALL·E omitted/null/url skips convert; GPT Image omitted encodes b64_json; custom omitted b64_json is an aio-proxy extension.
  • Edits accept official-max envelopes (357_564_416 JSON, 851_048_559 multipart). P1 has no lower default DoS cap; a future smaller ceiling is an explicit deployment extension.
  • Non-catalog Images Providers need a finite id set (models or preserved alias targets) including gpt-image-2.5-sunburst for the blank-model default. Providers configured only for gpt-image-2 must add the new model or clients must request the old model explicitly. A router.models metadata entry does not create a route.

Audio notes:

  • Raw Audio needs an openai-audio endpoint (or primary protocol).
  • Omitted model defaults to tts-1 for speech and whisper-1 for transcriptions/translations, so a non-catalog Audio Provider must list those ids (in models or as preserved alias targets) for the default to route.
  • POST /v1/audio/translations is raw passthrough only. Convert returns 501 unsupported_feature because the AI SDK transcription interface has no translation mode.
  • Convert also returns 501 unsupported_feature for stream_format, chunking_strategy, include, and stream.
  • An openai-audio Provider is probed with GET /v1/models, because a speech-only or transcription-only model rejects the other direction's request and would probe FAIL. A green probe means reachable with an accepted key, not that the model supports the direction you will call; a 401 is FAIL, and a gateway serving only /v1/audio/* shows FAIL in the Dashboard even when it works.
  • Audio usage is recorded only when upstream reports it. Duration is never converted to tokens.

Realtime notes:

  • Codex ChatGPT OAuth only. Official sessions, client_secrets, transcription sessions, translations, and SIP accept/reject/refer answer 501.
  • WebRTC media is not relayed; media stays client-to-upstream.
  • Sideband and hangup require the creating process's call pin; a restart forgets live calls.

Videos notes:

  • Raw Videos needs an openai-video endpoint (or primary protocol) and a finite id set including sora-2.
  • Omitted model defaults to sora-2.
  • Convert is not implemented (501 unsupported_feature).
  • Retrieve, content, delete, and remix require the creating process's pin; a restart forgets pins.
  • Edits and extensions accept JSON (video.id); official SDK multipart is 415.
  • GET /v1/videos and the character ports answer 501. /v1/videos/generations is not a proxy port.
  • Official Sora / Videos API shutdown is 2026-09-24; the /v1/videos wire remains for compatible gateways.

Remaining official Responses resource operations (GET /v1/responses/:id, DELETE /v1/responses/:id, POST /v1/responses/:id/cancel, GET /v1/responses/:id/input_items) return a protocol-shaped 501.

Call the OpenAI Responses endpoint:

curl http://127.0.0.1:9317/v1/responses \
  -H 'Content-Type: application/json' \
  -d '{"model":"gpt-5","input":"Introduce AIO Proxy in one sentence."}'

Dashboard and observability

The Dashboard is available at http://127.0.0.1:9317/dashboard. Use it to manage Providers and inspect runtime behavior:

  • Add, edit, and test Providers, including plugin OAuth login.
  • View request volume, token usage, and cost trends.
  • Search complete request traces and inspect the status and latency of each Provider attempt.

Set server.password to protect the Dashboard. It does not protect model API endpoints; use server.apiKeys for those. The login is kept in the browser, so new tabs and browser restarts do not ask for the password again; visiting at least once every 6 days renews it automatically. Changing server.password immediately invalidates the login on every device.

Network and security

Set the top-level proxy to configure a default HTTP(S) / SOCKS5 proxy. A Provider can inherit it, override it, or disable it with false. An api Provider can also set upstream request headers through headers.

SOCKS5 accepts socks://, socks5://, and socks5h:// URLs, optionally with percent-encoded username/password (for example, socks5://user:[email protected]:1080). All three use remote DNS; the default port is 1080. This applies to model requests, OAuth traffic, and plugin realtime connections.

Optional global proxy fallback is disabled by default:

{
  "proxy": "http://primary.example:8080",
  "proxyBackup": "socks5://backup.example:1080",
  "proxyFallback": true
}

Providers that omit proxy inherit this entire policy. A provider can set its own proxy, proxyBackup, and proxyFallback fields to replace the global policy entirely (even when its fallback is disabled); proxy: false uses a direct connection. Disabling proxyFallback preserves the backup address but only uses the primary. Enabling it requires both addresses. The dashboard exposes the backup address and switch in global settings and in each provider's independent proxy mode. Only choosing the inherit/global mode uses global settings. Each new proxy connection tries the primary first, with a 5-second proxy connection timeout before trying the backup. HTTP(S) proxies must support CONNECT for fallback. No upstream request is replayed, upstream HTTP errors do not trigger fallback, and failure of both proxies never falls back to direct access.

By default AIO Proxy binds to 127.0.0.1. Set server.host to another non-empty host (for example, 0.0.0.0) when clients need remote access. The proxy serves HTTP only, so terminate TLS with a reverse proxy, tunnel, or gateway before exposing it beyond a trusted network. Add server.apiKeys before doing so:

{
  "server": {
    "host": "0.0.0.0",
    "apiKeys": [{ "key": "{{env.AIO_PROXY_KEY}}", "label": "CI" }],
    "password": "a-dashboard-password",
  },
}

Each label is optional and only helps identify a key. With at least one key configured, every /v1/* and /v1beta/* request (including /v1/models) must send Authorization: Bearer <key> or X-API-Key: <key>; native Gemini clients may use X-Goog-Api-Key, ?key=, or ?auth_token=. Matched caller credentials are stripped before the request is forwarded upstream. An empty list leaves model APIs open. Set server.requireApiKey to false to keep the configured keys on file while accepting requests without one; the Dashboard exposes the same switch. Remote Dashboard access requires server.password and its Dashboard session. /admin/* remains loopback-only for local CLI control. Browser writes without a Dashboard password must come from a loopback Origin on the proxy port (127.0.0.1, localhost, [::1], or the configured loopback host). Direct loopback peers, including a local reverse proxy, are treated as local.

Agent integrations

aio-proxy supports two integration types: managed plugins for OpenCode, Pi, and oh-my-pi, and a Codex global configuration integration with two authentication modes. Plugin integrations install or update an adapter and keep the Agent's native login flow. The Codex integration edits the global config.toml; it does not install a plugin or replace native Codex login. Claude Code is connected the same way, through two keys in its global settings.json. Integrations are global to the current user and do not write project-local Agent config.

aio-proxy agent configure opencode
aio-proxy agent configure pi
aio-proxy agent configure omp
aiop agent configure codex
aiop agent configure claude-code
aio-proxy agent list --check
aio-proxy agent list --authorizations
aiop agent list --check
aio-proxy agent remove <target>
aiop agent remove codex

The Dashboard's Agents page shows the same state and, when opened on the machine running aio-proxy (not in a container), configures, repairs, and removes each Agent with one click, including the Codex setup and device approval.

Supported floors are OpenCode 1.17.10, Pi 0.84.2, and oh-my-pi 17.3.7. After configure, sign in with opencode auth login --provider aio-proxy or /login aio-proxy in Pi and oh-my-pi. Reload or restart the Agent so it loads the updated adapter. aio-proxy upgrade refreshes managed adapters the same way and also requires a reload.

When caller keys are enforced, set server.password so Device Approval can authorize the Agent. aio-proxy agent remove revokes the installation and deletes aio-proxy's managed files; it does not log the Agent out of its own host account. If the local control plane is offline, remove refuses and leaves files in place.

Claude Code configuration

aiop agent configure claude-code merges two keys into the env block of the global ~/.claude/settings.json, or of the directory selected by CLAUDE_CONFIG_DIR: ANTHROPIC_BASE_URL, set to the proxy's loopback address, and ANTHROPIC_AUTH_TOKEN. The token is required, because with a base URL alone Claude Code keeps using its saved claude.ai login. When server.apiKeys is empty, the command writes the non-secret aio-proxy-local placeholder and asks nothing. When keys exist, you choose one in the terminal or on the Dashboard's Agents page; it is checked against the proxy and stored as plain text in settings.json, which is then restricted to your user. A non-interactive run with keys configured fails without writing anything. Restart Claude Code for the change to take effect.

Every other setting and env entry is left as it is, and a symlinked or unparseable settings.json is refused. Ownership is recorded in ~/.claude/.aio-proxy/claude-code-config.json, outside the file Claude Code rewrites. aiop agent list reports the integration as managed, or as modified with the fields you edited, and --check probes the proxy with the configured token. aiop agent remove claude-code works offline and puts the two keys back to what they were, skipping any you have edited since; proxy keys are retained. Project-level .claude/settings.json files are never touched.

Codex configuration

aiop agent configure codex uses the global ~/.codex/config.toml, or the directory selected by CODEX_HOME, and takes effect after Codex is reopened. The default Provider ID is aio-proxy; the wizard accepts a custom Provider ID when it does not conflict with an existing entry. It preserves the top-level model and does not add a model selection step. The provider uses the proxy's Responses endpoint and displays AIO Proxy as its name.

The wizard defaults to keeping ChatGPT login available. In this mode it can use an existing proxy API Key; when no key exists, it skips key selection and uses the non-secret aio-proxy-local placeholder without creating a key or changing server.apiKeys. Native Codex login remains available. The command mode instead performs one AIO Proxy device approval during setup, stores the installation credential privately, and configures Codex to call aiop agent auth codex --installation-id <uuid>. On success, the helper's stdout contains only the raw bearer token followed by a newline; it emits no JSON, logs, device code, or refresh token. The helper returns a short-lived access token and refreshes it silently when needed. Command mode does not provide features that depend on ChatGPT login. Use aiop agent revoke <installation-id> to revoke command credentials.

The wizard can offer history migration after showing the source Provider IDs and counts. It filters selected legacy JSONL sessions by Provider ID and leaves unselected and unrelated history in place; this is not a deletion mechanism or a security isolation boundary. The verified compatibility experiment used codex-cli 0.146.0 and validated the observed configuration and app-server schema fields. The executable exposes no upstream revision, and compatibility beyond that observed version and format is unverified. Native and paginated state migration remains blocked pending an independently verified storage contract.

If a configuration operation is pending, rerun aiop agent configure codex to enter the recovery prompt. A completed or partial history migration can be restored with aiop agent configure codex --restore-migration <operation-id>. Journals and JSONL backups are kept under $CODEX_HOME/.aio-proxy/migrations/<operation-id>/ and its backups/ directory. aiop agent list --check is read-only and reports the configured Provider ID, active Provider ID, configuration status, and proxy connection status; aiop agent list --authorizations shows command installations. aiop agent remove codex restores aio-proxy-owned fields while preserving unrelated settings and later user edits, retains proxy Keys, and does not reverse-migrate history. Command removal revokes the installation before deleting local credentials; if the control plane is offline, it leaves a retryable state. Native and paginated state migration, Computer Use, and plugin compatibility are not implied by this integration.

aio-proxy does not copy an upstream API key or a shared embedded SK into an Agent.

Common commands

aio-proxy status --deep
aio-proxy provider list --probe
aio-proxy doctor
aio-proxy --help

Contributing

See the contribution guide for development setup and submission guidelines.

License

MIT