@molecule/api-ai-local
v1.0.5
Published
AIProvider bond for any OpenAI-compatible local inference server (Ollama, LM Studio, llama.cpp, vLLM) — configurable base URL, keyless by default, streamed via raw fetch.
Readme
@molecule/api-ai-local
Auto-generated, AI-first package reference for the molecule.dev ecosystem. It is written to be read by coding agents as much as by people, and is generated from this package's source — edit
src/index.tsJSDoc, not this file.
Local (OpenAI-compatible) ai-local provider for molecule.dev.
Streams chat completions from any local inference server that speaks the
OpenAI chat/completions protocol (Ollama, LM Studio, llama.cpp, vLLM),
keyless by default.
Quick Start
// npm install @molecule/api-ai-local --workspace=api
import { setProvider, requireProvider } from '@molecule/api-ai'
import { provider } from '@molecule/api-ai-local'
// Named registration — several AI providers can be bonded side by side
// (getProviderByName('local') targets this one); the FIRST one
// registered also answers requireProvider() (see @molecule/api-ai).
setProvider('local', provider) // keyless: LOCAL_AI_BASE_URL (Ollama http://localhost:11434/v1 by default), LOCAL_AI_MODEL (llama3.1)
let reply = ''
for await (const event of requireProvider().chat({
messages: [{ role: 'user', content: 'Hello!' }],
})) {
if (event.type === 'text') reply += event.content // or forward the chunk to the client (SSE)
}Type
provider
Installation
npm install @molecule/api-ai-local @molecule/api-ai @molecule/api-bond @molecule/api-i18n @molecule/api-secretsAPI
Interfaces
LocalConfig
Local (OpenAI-compatible) provider configuration.
interface LocalConfig {
/** Called on each rate-limited/busy upstream response, before any retry sleep. */
onRateLimit?: AiRateLimitCallback
/**
* Base URL of the OpenAI-compatible endpoint, INCLUDING the version segment
* (e.g. `http://localhost:11434/v1`). Overrides `LOCAL_AI_BASE_URL` /
* `OLLAMA_BASE_URL`. Defaults to Ollama's `http://localhost:11434/v1`.
*/
baseUrl?: string
/**
* Optional API key. Most local servers ignore auth; when omitted (and no
* `LOCAL_AI_API_KEY` env var is set) no `Authorization` header is sent.
*/
apiKey?: string
/**
* Default model when a call doesn't specify one. Overrides `LOCAL_AI_MODEL`.
* Defaults to `llama3.1`.
*/
model?: string
}Classes
LocalAIProvider
Local (OpenAI-compatible) chat provider implementing the AIProvider
interface. Mirrors @molecule/api-ai-openai so the same handler code can
dispatch to a local endpoint.
Functions
createProvider(config)
Create a local (OpenAI-compatible) AI provider instance.
Constructs WITHOUT requiring any secret — local endpoints run keyless.
function createProvider(config?: LocalConfig): AIProviderconfig— Local provider configuration.
Returns: An AIProvider backed by an OpenAI-compatible local endpoint.
Constants
aiLocalSecretDefinitions
Optional secret/config overrides recognised by the local AI bond.
const aiLocalSecretDefinitions: SecretDefinition[]provider
The provider implementation. Constructs keyless — no secret required.
const provider: AIProviderCore Interface
Implements @molecule/api-ai interface.
Bond Wiring
Setup function to register this provider with the bond system:
import { bond } from '@molecule/api-bond'
import { provider } from '@molecule/api-ai-local'
export function setupAiLocal(): void {
bond('ai', 'local', provider)
}Injection Notes
Requirements
Peer dependencies:
@molecule/api-ai^1.0.1@molecule/api-bond^1.0.1@molecule/api-i18n^1.0.1@molecule/api-secrets^1.0.1
Runtime Dependencies
@molecule/api-ai@molecule/api-bond@molecule/api-i18n@molecule/api-secrets
Error message disambiguation: a plain 400 that ISN'T a context-length error (bad param, malformed tool schema) gets its own non-retryable message distinct from the generic "AI service error. Please try again." used for retryable failures.
E2E Tests
Integration checklist — drive the real UI (live preview, no mocks), adapt each item to this app's actual chat/AI screens, and check every box off one by one. A box you can't check is an integration bug to fix — not a skip. The sandbox HAS an AI provider bonded, so the flow runs live end-to-end; AI output is NON-DETERMINISTIC, so assert on STRUCTURE/behavior, not exact text:
- [ ] A message sent through the real chat UI comes back as a RELEVANT AI reply — not an echo of the prompt, a hardcoded stub, or an empty bubble. Ask something with a checkable answer (e.g. "What is 2 + 2?") and confirm the response actually contains it ("4"), proving a live model answered.
- [ ] If the app streams, tokens render INCREMENTALLY — text grows word by
word in the UI, not one final blob dumped after a long frozen spinner. (A
streamed
chat()yieldstextchunks then a finaldone; a single late blob means the reply was awaited whole and streaming is broken.) - [ ] Multi-turn CONTEXT is preserved: a follow-up that refers back to the
previous turn (e.g. after "2 + 2", ask "now double that" -> understood as 8) works — proving the full
messageshistory is sent, not just the last line. - [ ] A provider failure (bad/missing key, rate limit, timeout) surfaces as a graceful in-UI error message, NOT a crash, blank screen, a spinner that never resolves, or an unhandled 500. Force one and watch the UI recover.
- [ ] The provider key + the provider call are SERVER-side only: the key
never reaches the browser (check the network tab, the JS bundle, and page
globals), and no route proxies arbitrary prompts to the model without auth
- a token cap — an open AI endpoint is an unbounded bill and abuse vector.
