@corbits/ollama-adapter
v0.2.0
Published
Ollama as an Interchange inference provider: OpenAI-compatible and Anthropic messages factories, one reasoning setting, think-tag stripping, inline tool JSON repair.
Readme
@corbits/ollama-adapter
@intx/inference provider adapters for Ollama's OpenAI-compatible /v1/chat/completions and Anthropic-compatible /v1/messages surfaces. A Corbits inference provider that plugs into the Interchange adapter registry and also works in any host that runs @intx/inference.
Why @corbits/ollama-adapter?
- Clean streams from local models. Many Ollama models emit chain-of-thought inside
<think>tags or tool calls as raw JSON in the text.createOllamaAdapterreclassifies both into thinking and tool-call events before they reach the agent. - One reasoning setting for both surfaces.
reasoningis sent asreasoning_efforton/v1/chat/completionsand as the Anthropicthinkingswitch on/v1/messages, set per model or as a default. - Local or Ollama Cloud. Point the same source at a local
ollama serveor at Ollama Cloud. Both surfaces keep their native request shape and useAuthorization: Bearer.
It does not cover Ollama's native /api endpoints.
Install
bun add @corbits/ollama-adapter @intx/inference@^0.4.0 @intx/log@^0.4.0 @intx/types@^0.4.0Runs on Bun >= 1.2 or Node >= 24.
Quickstart
Needs Ollama running locally with the model pulled (ollama pull gpt-oss:20b).
import { createDependencies, runInference } from "@intx/inference";
import { createOllamaAdapter } from "@corbits/ollama-adapter";
const deps = createDependencies({
has: (provider) => provider === "ollama",
resolve: (source, quirks) => createOllamaAdapter(source, quirks),
});
let seq = 0;
for await (const event of runInference({
deps,
source: {
id: "ollama",
provider: "ollama",
baseURL: "http://localhost:11434/v1",
credentialId: "ollama",
model: "gpt-oss:20b",
quirks: { default: { reasoning: "low" } },
},
turns: [
{
role: "user",
timestamp: Date.now(),
content: [{ type: "text", text: "Say hello." }],
},
],
nextSeq: () => seq++,
// Local Ollama ignores the key; Ollama Cloud needs a real one.
readMaterial: () => ({ secret: "ollama" }),
})) {
if (event.type === "inference.text.delta")
process.stdout.write(event.data.token);
if (event.type === "inference.error")
throw new Error(event.data.error.message);
}
process.stdout.write("\n");Where it fits
Interchange runs AI agents as principals (accounts that hold their own identity, permissions and credentials). Corbits packages add what an agent product needs around it.
- Runs in: the agent sidecar (the runtime next to each agent), or any process that calls
runInference. No hub (the multi-tenant control plane) is required. - Plugs into: the
@intx/inferenceadapter registry, as the factory for an Ollama provider id. - Pairs with:
@corbits/openai-responsesand@corbits/system-one, the other Corbits inference providers.
Reference
| Export | Description |
| ------------------------------ | ------------------------------------------------------------------------------------- |
| createOllamaAdapter | AdapterFactory for /v1/chat/completions. Applies every config field. |
| createOllamaAnthropicAdapter | AdapterFactory for /v1/messages. Accepts the same config and applies reasoning. |
| OllamaAdapterConfig | Schema and type for the source's quirks: { default?, perModel? } of overrides. |
| OllamaAdapterOverride | Schema and type for one override. Unknown keys are rejected. |
| Reasoning | Schema and type for reasoning: boolean \| "low" \| "medium" \| "high" \| "max". |
Set the source's baseURL to http://localhost:11434/v1 or https://ollama.com/v1. The factories append /chat/completions and /messages. Send images as base64.
Overrides
A perModel entry wins field by field over default. An unset field leaves the request unchanged.
| Field | Type | createOllamaAdapter sends | createOllamaAnthropicAdapter sends |
| ----------------- | ------------------- | --------------------------- | ------------------------------------ |
| maxOutputTokens | positive integer | max_tokens | nothing |
| reasoning | see the table below | reasoning_effort | thinking |
| reasoning | createOllamaAdapter | createOllamaAnthropicAdapter |
| ----------------------------------------- | ---------------------------- | ---------------------------------------------------- |
| false | reasoning_effort: "none" | thinking: { type: "disabled" } |
| true | reasoning_effort: "medium" | thinking: { type: "enabled", budget_tokens: 1024 } |
| "low" / "medium" / "high" / "max" | reasoning_effort as given | thinking: { type: "enabled", budget_tokens: 1024 } |
/v1/messages has no effort levels, so every level enables thinking with the same budget. On that factory, reasoning: false always disables thinking; otherwise a per-call thinking option wins. With reasoning set, a thinking budget_tokens at or above max_tokens throws at buildRequest.
reasoning_effort is sent only when reasoning is set. default.reasoning applies to every model, and Ollama rejects reasoning_effort with a 400 for models that cannot reason, so in a mixed fleet set reasoning per model.
Ollama ignores num_ctx on /v1. Set the context length on the server with OLLAMA_CONTEXT_LENGTH=32768 ollama serve, or with PARAMETER num_ctx 32768 in the model's Modelfile.
Using with Interchange
The sidecar loads custom adapters from the SIDECAR_ADAPTER_MANIFEST env var, a JSON array of { provider, specifier, export } entries. Install this package in the sidecar's workspace and register a provider id:
SIDECAR_ADAPTER_MANIFEST='[
{"provider":"ollama","specifier":"@corbits/ollama-adapter","export":"createOllamaAdapter"}
]'Use createOllamaAnthropicAdapter as the export for the /v1/messages surface. A hub that spawns sidecars forwards its own SIDECAR_ADAPTER_MANIFEST to each one, so set it once on the hub. Sources with that provider then resolve to this adapter, and each source's quirks holds its OllamaAdapterConfig.
Upgrading from 0.1
reasoningEffortandthinkare replaced byreasoning. Old configs still load, and each key logs one warning per config location.reasoningEffortis read asreasoningunlessreasoningis set.thinkis dropped, since Ollama ignored it on/v1; setreasoningto turn reasoning on.numCtxis removed because Ollama ignores it on/v1. Old configs still load: the key logs one warning and is dropped. Set the context length on the server instead.createOllamaAnthropicAdapternow appliesreasoningand throws when the thinking budget is at or abovemax_tokens.parseOllamaAdapterConfig,resolveOverrideand the think-tag and inline-JSON helpers are no longer exported.@intx/logis a new peer. All@intx/*peers are now^0.4.0, which excludes 0.5.
