@intentqa/llm-ollama
v2.0.0
Published
Ollama planner adapter for the intent-driven QA SDK. Local models, no API key.
Readme
@intentqa/llm-ollama
Ollama planner adapter for intentqa. Local models, no API key, no bill.
pnpm add -D @intentqa/llm-ollama
ollama serve
ollama pull qwen3:8b
npx intentqa generate --intent "a premium user can apply a coupon"Read this before you rely on it
A small local model writes materially worse plans than a frontier model. It picks the wrong catalog entry more often, misses scenarios a reviewer would expect, invents step parameters, and fails schema validation outright often enough to notice. That is a property of the models, not of this adapter, and no amount of prompt work here fixes it.
What this package is actually for: a zero-cost, offline, no-key loop while you are
iterating on a vocabulary. Draft a catalog, throw ten intents at it, and see which ones
have no reasonable plan — that tells you which steps are missing, and it tells you that
just as well from a mediocre model as from a good one, for free and on a plane. When the
vocabulary settles and the plan itself is the deliverable, switch to
@intentqa/llm-anthropic, @intentqa/llm-openai or @intentqa/llm-google. The prompt and
the plan schema live in intentqa, so that switch changes nothing but the provider.
No dependencies beyond intentqa itself. This adapter posts to /api/chat with the
global fetch and hand-rolled JSON rather than taking ollama as a dependency — the same
call @intentqa/mcp makes with JSON-RPC. One request shape does not justify putting a
vendor SDK, and its transitive tree, into everyone's lockfile.
Options
Defaults to qwen3:8b against http://127.0.0.1:11434. Every default is an option:
import { ollamaPlanner } from "@intentqa/llm-ollama";
export const planner = ollamaPlanner({
model: "qwen3:30b", // any tag you have pulled
host: "http://gpu-box.lan:11434",
maxTokens: 24000,
temperature: 0.2,
});host also reads OLLAMA_HOST, bare (gpu-box.lan:11434) or with a scheme. baseURL is
accepted as an alias for host so a config can be moved between adapters unchanged; if
both are given, host wins. apiKey is accepted and unused unless you have fronted
Ollama with a proxy that wants a bearer token.
The plan schema is passed straight through as Ollama's format, which hands it to
llama.cpp's grammar-constrained sampling. Constrained decoding makes the reply parse; it
does not make the reply right. The compiler still checks every step id against the
catalog, and the reviewer still reads the diff.
