npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@earendil-works/pi-ai

v0.99.1

Published

Unified LLM API with automatic model discovery and provider configuration

Readme

@earendil-works/pi-ai

Unified LLM API with provider collections, automatic auth resolution, token and cost tracking, and simple context persistence and hand-off to other models mid-session.

Note: The chat catalog only includes models that support tool calling (function calling), as this is essential for agentic workflows. Image and classifier catalogs use their operation-specific capabilities.

Table of Contents

Supported Providers

  • OpenAI
  • Ant Ling
  • Azure OpenAI (Responses)
  • OpenAI Codex (legacy) (ChatGPT Plus/Pro subscription, requires OAuth, see below)
  • Radius (API key or OAuth, with a dynamically refreshed gateway catalog)
  • TypeSafe (System One classifier API)
  • DeepSeek
  • NVIDIA NIM
  • Anthropic
  • Google
  • Vertex AI (Gemini via Vertex AI)
  • Mistral
  • Groq
  • Cerebras
  • Cloudflare AI Gateway
  • Cloudflare Workers AI
  • xAI
  • OpenRouter
  • Vercel AI Gateway
  • ZAI Coding Plan (Global) (with separate China provider)
  • MiniMax (with separate China provider)
  • Together AI
  • Baseten
  • Hugging Face
  • Moonshot AI (with separate China provider)
  • GitHub Copilot (requires OAuth, see below)
  • Amazon Bedrock
  • OpenCode Zen
  • OpenCode Go
  • Fireworks (uses OpenAI- and Anthropic-compatible APIs)
  • Kimi For Coding (Moonshot AI subscription endpoint, uses Anthropic-compatible API)
  • Meta (Model API, uses OpenAI Responses-compatible API)
  • Qwen Token Plan (separate Individual and existing catalogs, with a separate China provider)
  • Xiaomi MiMo (defaults to API billing endpoint, with separate Token Plan providers for cn/ams/sgp regions)
  • Any OpenAI-compatible API: Ollama, vLLM, LM Studio, etc.

Installation

npm install @earendil-works/pi-ai

TypeBox exports are re-exported from @earendil-works/pi-ai: Type, Static, and TSchema.

Quick Start

You build a Models collection of providers and stream through it. The quickest start registers every built-in provider; apps that care about bundle size register individual providers instead (see Provider Factories and Bundling and Tree Shaking).

import { Type, type Context, type Tool } from '@earendil-works/pi-ai';
import { builtinModels } from '@earendil-works/pi-ai/providers/all';

// A Models collection with every built-in provider registered
const models = builtinModels();

// Sync lookup against the collection
const model = models.getModel('openai', 'gpt-4o-mini')!;

// Define tools with TypeBox schemas for type safety and validation
const tools: Tool[] = [{
  name: 'get_time',
  description: 'Get the current time',
  parameters: Type.Object({
    timezone: Type.Optional(Type.String({ description: 'Optional timezone (e.g., America/New_York)' }))
  })
}];

// Build a conversation context (easily serializable and transferable between models)
const context: Context = {
  systemPrompt: 'You are a helpful assistant.',
  messages: [{ role: 'user', content: 'What time is it?', timestamp: Date.now() }],
  tools
};

// Option 1: Streaming with all event types.
// Auth resolves through the provider (OPENAI_API_KEY from the environment here).
const s = models.stream(model, context);

for await (const event of s) {
  switch (event.type) {
    case 'start':
      console.log(`Starting with ${event.partial.model}`);
      break;
    case 'text_start':
      console.log('\n[Text started]');
      break;
    case 'text_delta':
      process.stdout.write(event.delta);
      break;
    case 'text_end':
      console.log('\n[Text ended]');
      break;
    case 'thinking_start':
      console.log('[Model is thinking...]');
      break;
    case 'thinking_delta':
      process.stdout.write(event.delta);
      break;
    case 'thinking_end':
      console.log('[Thinking complete]');
      break;
    case 'toolcall_start':
      console.log(`\n[Tool call started: index ${event.contentIndex}]`);
      break;
    case 'toolcall_delta':
      // Partial tool arguments are being streamed
      const partialCall = event.partial.content[event.contentIndex];
      if (partialCall.type === 'toolCall') {
        console.log(`[Streaming args for ${partialCall.name}]`);
      }
      break;
    case 'toolcall_end':
      console.log(`\nTool called: ${event.toolCall.name}`);
      console.log(`Arguments: ${JSON.stringify(event.toolCall.arguments)}`);
      break;
    case 'done':
      console.log(`\nFinished: ${event.reason}`);
      break;
    case 'error':
      console.error(`Error: ${event.error.errorMessage}`);
      break;
  }
}

// Get the final message after streaming, add it to the context
const finalMessage = await s.result();
context.messages.push(finalMessage);

// Handle tool calls if any
const toolCalls = finalMessage.content.filter(b => b.type === 'toolCall');
for (const call of toolCalls) {
  const result = call.name === 'get_time'
    ? new Date().toLocaleString('en-US', {
        timeZone: call.arguments.timezone || 'UTC',
        dateStyle: 'full',
        timeStyle: 'long'
      })
    : 'Unknown tool';

  // Add tool result to context (supports text and images)
  context.messages.push({
    role: 'toolResult',
    toolCallId: call.id,
    toolName: call.name,
    content: [{ type: 'text', text: result }],
    isError: false,
    timestamp: Date.now()
  });
}

// Continue if there were tool calls
if (toolCalls.length > 0) {
  const continuation = await models.complete(model, context);
  context.messages.push(continuation);
  console.log('After tool execution:', continuation.content);
}

console.log(`Total tokens: ${finalMessage.usage.input} in, ${finalMessage.usage.output} out`);
console.log(`Cost: $${finalMessage.usage.cost.total.toFixed(4)}`);

// Option 2: Get complete response without streaming
const response = await models.complete(model, context);

for (const block of response.content) {
  if (block.type === 'text') {
    console.log(block.text);
  } else if (block.type === 'toolCall') {
    console.log(`Tool: ${block.name}(${JSON.stringify(block.arguments)})`);
  }
}

Snippets in the rest of this README assume a models collection set up like this (with the relevant providers registered).

Providers and Models

A provider is the runtime unit: it owns its model catalog, its auth (API key resolution, OAuth flows), and its stream behavior. A Models collection holds providers and routes every request to the provider that owns the model.

Providers internally share API implementations (the wire protocols): Anthropic models use anthropic-messages, OpenAI uses openai-responses, while xAI, Groq, Cerebras, OpenRouter, and most others share openai-completions. Mixed-API providers (GitHub Copilot, OpenCode Zen) dispatch per model.

Provider Factories

For apps that only need specific providers, there is one factory per built-in provider, each a subpath import that pulls only that provider's catalog:

import { anthropicProvider } from '@earendil-works/pi-ai/providers/anthropic';
import { openaiProvider } from '@earendil-works/pi-ai/providers/openai';
import { openrouterProvider } from '@earendil-works/pi-ai/providers/openrouter';
import { amazonBedrockProvider } from '@earendil-works/pi-ai/providers/amazon-bedrock';
// ...one module per provider in the Supported Providers list

const models = createModels();
models.setProvider(anthropicProvider());
models.setProvider(openrouterProvider());

Provider factories import their model catalog and a lazy API wrapper. They do not import other providers. With bundler code splitting, SDK implementations (@anthropic-ai/sdk, openai, @google/genai, etc.) stay in lazy chunks loaded on the first request to a model of that API.

All Built-in Providers

For apps that want everything (as in Quick Start):

import { builtinModels } from '@earendil-works/pi-ai/providers/all';

const models = builtinModels(); // a Models collection with every built-in provider registered

This imports all catalogs and every built-in provider factory. It is the heavy, explicit entrypoint. builtinModels() accepts the same options as createModels() (credentials, authContext); builtinProviders() returns the provider array if you want to register them on your own collection.

Querying Models

Reads are synchronous and return the last-known lists:

const providers = models.getProviders();           // registered Provider objects
const provider = models.getProvider('anthropic');  // one provider

const all = models.getModels();                    // every chat model across providers
const anthropicModels = models.getModels('anthropic');
const model = models.getModel('anthropic', 'claude-sonnet-4-5');

for (const m of anthropicModels) {
  console.log(`${m.id}: ${m.name}`);
  console.log(`  API: ${m.api}`);
  console.log(`  Context: ${m.contextWindow} tokens`);
  console.log(`  Vision: ${m.input.includes('image')}`);
  console.log(`  Reasoning: ${m.reasoning}`);
}

The unqualified reads getModels()/getModel()/getAvailable() return chat models (Model<Api>) usable with stream(). The *OfType reads return one model type, and getAllModels()/getAllAvailable() return every type as AnyModel:

const images = models.getModelsOfType('image', 'openrouter');           // ImageModel[]
const flux = models.getModelOfType('image', 'openrouter', 'black-forest-labs/flux.2-pro');
const jev = models.getModelOfType('classifier', 'typesafe', 'jev-latest');
const availableImages = await models.getAvailableOfType('image');
const everything = models.getAllModels();                               // AnyModel[]

The model's type decides which operation accepts it: chat models stream, type: "image" models generate images, and type: "classifier" models classify structured state. type is optional on chat models, so a model without type is a chat model. Do not compare type directly; narrow mixed lists with isModelType() or read the effective type with getModelType():

import { isModelType } from '@earendil-works/pi-ai';

for (const model of models.getAllModels()) {
  if (isModelType(model, 'image')) {
    // model: ImageModel<ImageApi>
  }
}

IDs are unique within each provider and type; one upstream model may have separate entries for different operations. On a provider, getModels() returns chat models and the optional getAllModels() returns every type; providers with only chat models can omit it.

Dynamically listed chat models are typed Model<Api>. Narrow with the hasApi() guard when you need API-specific option typing:

import { hasApi } from '@earendil-works/pi-ai';

const m = models.getModel('anthropic', 'claude-sonnet-4-5');
if (m && hasApi(m, 'anthropic-messages')) {
  // m: Model<'anthropic-messages'> — stream options fully typed
  models.stream(m, context, { thinkingEnabled: true, thinkingBudgetTokens: 2048 });
}

Static Catalog Reads

For tooling that wants the generated built-in catalog with full literal typing (provider and model IDs auto-complete), independent of any collection:

import {
  getAllBuiltinModels,
  getBuiltinClassifierModel,
  getBuiltinClassifierModels,
  getBuiltinImageModel,
  getBuiltinImageModels,
  getBuiltinModel,
  getBuiltinModels,
  getBuiltinProviders,
} from '@earendil-works/pi-ai/providers/all';

const model = getBuiltinModel('openai', 'gpt-4o-mini'); // typed Model<'openai-responses'>
const radius = getBuiltinModel('radius', 'balanced');   // typed Model<'pi-messages'>
const flux = getBuiltinImageModel('openrouter', 'black-forest-labs/flux.2-pro');
const jev = getBuiltinClassifierModel('typesafe', 'jev-latest');
const providers = getBuiltinProviders();
const openrouterChat = getBuiltinModels('openrouter');        // Model[]
const openrouterImages = getBuiltinImageModels('openrouter'); // ImageModel[]
const typesafeClassifiers = getBuiltinClassifierModels('typesafe'); // ClassifierModel[]
const openrouterAll = getAllBuiltinModels('openrouter');      // AnyModel[]

Dynamic Providers

Providers may have dynamic model lists (a llama.cpp server, a live OpenRouter listing). Reads stay sync; fetching is an explicit async verb:

// getModels() returns the last-known list (empty before the first refresh)
await models.refresh({ providers: ['llamacpp'] }); // refresh one provider
await models.refresh();                            // refresh all providers concurrently, best-effort
const fresh = models.getModel('llamacpp', 'qwen3-30b');

Static built-in providers are no-ops for refresh(). Radius is both static and dynamic: it ships the public radius.pi.dev catalog for synchronous API lookup, then overlays cached and freshly fetched /v1/config models when refreshed with configured auth. See createProvider() for building a dynamic provider.

Auth

Every provider owns its auth: how API keys resolve (stored credentials, environment variables, ambient sources like AWS profiles or gcloud ADC) and, where supported, OAuth login/refresh flows.

How Auth Resolves

When you call models.stream(), the collection resolves auth through the owning provider and merges it into the request. Explicit per-request values always win:

// Resolved through the provider (env var, stored credential, OAuth token):
await models.complete(model, context);

// Explicit key wins over anything the provider would resolve:
await models.complete(model, context, { apiKey: 'sk-explicit' });

You can inspect resolution without making a request. Pass a provider ID for provider-scoped auth, or a model to include its static model.headers:

const providerAuth = await models.getAuth(model.provider);
const modelAuth = await models.getAuth(model);

if (modelAuth) {
  console.log(`configured via ${modelAuth.source}`); // e.g. "ANTHROPIC_API_KEY", "OAuth", "stored credential"
  console.log(modelAuth.auth.headers);              // Provider auth headers + model.headers
} else {
  console.log('not configured');
}

Both overloads resolve credentials, refresh expired OAuth when necessary, and may return an auth-derived apiKey, headers, or baseUrl. getAuth() resolves undefined for unconfigured providers and rejects with ModelsError when something is actually broken ("oauth": token refresh failed, credential preserved for re-login; "auth": key resolution or credential store failure). Request paths surface the same failures as stream errors.

getAuth(), checkAuth(), getAvailable(), login, and logout accept optional caller cancellation through their existing options or interaction objects and remain unbounded when no signal is supplied. Provider login, ApiKeyAuth.check, ApiKeyAuth.resolve, and OAuthAuth.refresh implementations always receive a concrete signal and must honor it for blocking work.

Transforming Request Headers

Models.stream(), complete(), streamSimple(), and completeSimple() accept a Models-only transformHeaders option. It runs once after provider auth, model.headers, and explicit options.headers have been merged, but before provider dispatch:

const response = await models.completeSimple(model, context, {
  headers: { "X-Client": "my-app" },
  transformHeaders: async (headers) => ({
    ...headers,
    "X-Request-ID": crypto.randomUUID(),
  }),
});

The ordering is:

provider auth headers -> model.headers -> explicit options.headers -> transformHeaders -> Provider.stream*()

Header names are merged case-insensitively. Explicit headers override auth/model headers, and the transform has final control; returning null for a header suppresses lower-level defaults that support deletion.

transformHeaders belongs to Models, not Provider. A Models implementation must consume it and remove it before calling Provider.stream*(). Provider implementations continue receiving ordinary ApiStreamOptions or SimpleStreamOptions and never handle the transform themselves. Use this option instead of calling getAuth(model) before stream*(), which would resolve request auth twice.

Credential Store

Stored credentials (API keys entered interactively, OAuth tokens) live in a CredentialStore — one type-tagged credential per provider. pi-ai ships an in-memory default; apps inject persistent storage:

import { createModels, type CredentialStore } from '@earendil-works/pi-ai';

const models = createModels({ credentials: myFileBackedStore });
// builtinModels() takes the same options:
// const models = builtinModels({ credentials: myFileBackedStore });

The contract is small: read(providerId), list() for non-secret { providerId, type } metadata, modify(providerId, fn) (the only write path — a serialized read-modify-write), and delete(providerId). Each operation accepts optional cancellation options. Enumeration must not resolve secrets or execute configured key commands. OAuth token refresh runs inside modify, so concurrent requests and processes cannot double-refresh a rotated token. A stored credential owns its provider: environment variables are only consulted when nothing is stored, and a failed refresh never silently falls back to an env key.

API-key credentials use the same discriminator as pi's auth.json and can carry provider-scoped env/config values:

const credential = {
  type: 'api_key',
  key: '...',
  env: {
    CLOUDFLARE_ACCOUNT_ID: 'account-id',
    CLOUDFLARE_GATEWAY_ID: 'gateway-id'
  }
} as const;

Environment Variables

Built-in providers resolve these env vars (Node.js; in browsers pass apiKey explicitly):

| Provider | Environment Variable(s) | |----------|------------------------| | OpenAI | OPENAI_API_KEY | | Ant Ling | ANT_LING_API_KEY | | Azure OpenAI | AZURE_OPENAI_API_KEY + AZURE_OPENAI_BASE_URL (e.g. https://{resource}.ai.azure.com) or AZURE_OPENAI_RESOURCE_NAME. Supports *.openai.azure.com, *.cognitiveservices.azure.com and *.ai.azure.com; root endpoints auto-normalize to /openai/v1. Optional: AZURE_OPENAI_API_VERSION (default v1), AZURE_OPENAI_DEPLOYMENT_NAME_MAP. | | Anthropic | ANTHROPIC_API_KEY or ANTHROPIC_OAUTH_TOKEN | | Radius | RADIUS_API_KEY | | TypeSafe | TYPESAFE_API_KEY | | DeepSeek | DEEPSEEK_API_KEY | | NVIDIA NIM | NVIDIA_API_KEY | | Google | GEMINI_API_KEY | | Vertex AI | GOOGLE_CLOUD_API_KEY or GOOGLE_CLOUD_PROJECT (or GCLOUD_PROJECT) + GOOGLE_CLOUD_LOCATION + ADC | | Mistral | MISTRAL_API_KEY | | Groq | GROQ_API_KEY | | Cerebras | CEREBRAS_API_KEY | | Cloudflare AI Gateway | CLOUDFLARE_API_KEY + CLOUDFLARE_ACCOUNT_ID + CLOUDFLARE_GATEWAY_ID | | Cloudflare Workers AI | CLOUDFLARE_API_KEY + CLOUDFLARE_ACCOUNT_ID | | xAI | XAI_API_KEY | | Fireworks | FIREWORKS_API_KEY | | Together AI | TOGETHER_API_KEY | | Baseten | BASETEN_API_KEY | | OpenRouter | OPENROUTER_API_KEY | | Vercel AI Gateway | AI_GATEWAY_API_KEY | | ZAI Coding Plan (Global) | ZAI_API_KEY | | ZAI Coding Plan (China) | ZAI_CODING_CN_API_KEY | | MiniMax (Global) | MINIMAX_API_KEY | | MiniMax (China) | MINIMAX_CN_API_KEY | | Moonshot AI / Moonshot AI (China) | MOONSHOT_API_KEY | | Hugging Face | HF_TOKEN | | OpenCode Zen / OpenCode Go | OPENCODE_API_KEY | | Kimi For Coding | KIMI_API_KEY | | Meta | META_API_KEY | | Qwen Token Plan (existing catalog) | QWEN_TOKEN_PLAN_API_KEY | | Qwen Token Plan (Individual) | QWEN_TOKEN_PLAN_API_KEY | | Qwen Token Plan (China) | QWEN_TOKEN_PLAN_CN_API_KEY | | Xiaomi MiMo (API billing) | XIAOMI_API_KEY | | Xiaomi MiMo Token Plan (China) | XIAOMI_TOKEN_PLAN_CN_API_KEY | | Xiaomi MiMo Token Plan (Amsterdam) | XIAOMI_TOKEN_PLAN_AMS_API_KEY | | Xiaomi MiMo Token Plan (Singapore) | XIAOMI_TOKEN_PLAN_SGP_API_KEY | | GitHub Copilot | COPILOT_GITHUB_TOKEN |

qwen-token-plan-individual and qwen-token-plan share the international endpoint and QWEN_TOKEN_PLAN_API_KEY. The Individual provider exposes only the models documented for Individual subscriptions, while the existing provider retains its broader catalog for backward compatibility. Stored credentials remain provider-scoped, so save the key under the provider ID you register.

Amazon Bedrock resolves ambient AWS credentials (AWS_PROFILE, access key pairs, AWS_BEARER_TOKEN_BEDROCK, ECS task roles, web identity tokens); its provider-owned login flow supports bearer tokens, AWS profiles, and the existing credential chain. Vertex AI resolves either an explicit key or gcloud Application Default Credentials plus project/location, with a provider-owned login flow for API keys, ADC, and service-account files.

Tools

Tools enable LLMs to interact with external systems. This library uses TypeBox schemas for type-safe tool definitions with automatic validation using TypeBox's built-in validator and value conversion utilities. TypeBox schemas can be serialized and deserialized as plain JSON, making them ideal for distributed systems.

Defining Tools

import { Type, type Tool, StringEnum } from '@earendil-works/pi-ai';

// Define tool parameters with TypeBox
const weatherTool: Tool = {
  name: 'get_weather',
  description: 'Get current weather for a location',
  parameters: Type.Object({
    location: Type.String({ description: 'City name or coordinates' }),
    units: StringEnum(['celsius', 'fahrenheit'], { default: 'celsius' })
  })
};

// Note: For Google API compatibility, use StringEnum helper instead of Type.Enum
// Type.Enum generates anyOf/const patterns that Google doesn't support

const bookMeetingTool: Tool = {
  name: 'book_meeting',
  description: 'Schedule a meeting',
  parameters: Type.Object({
    title: Type.String({ minLength: 1 }),
    startTime: Type.String({ format: 'date-time' }),
    endTime: Type.String({ format: 'date-time' }),
    attendees: Type.Array(Type.String({ format: 'email' }), { minItems: 1 })
  })
};

Constrained Sampling for Tools

Tools can opt in to provider-side constrained sampling. For JSON-schema tools, strict: 'prefer' uses provider-side strict schema enforcement when supported and otherwise falls back to normal tool calling. strict: 'require' fails the request when the active provider/model cannot honor it. Set constrainedSampling: false to explicitly opt out; it behaves the same as omitting the field.

const strictTool: Tool = {
  name: 'edit_file',
  description: 'Edit a file',
  parameters: Type.Object({
    path: Type.String(),
    content: Type.String()
  }, { additionalProperties: false }),
  constrainedSampling: { type: 'json_schema', strict: 'prefer' }
};

Strict JSON-schema constrained sampling is supported for OpenAI, Anthropic, supported Amazon Bedrock Converse models, Mistral, and Gemini 3 tool calls through the Google Generative AI and Vertex adapters. Google uses VALIDATED function-calling mode (or ANY when explicitly requested); earlier Gemini versions fall back for strict: 'prefer' and reject strict: 'require' because they do not enforce required parameters. Bedrock strict-tool capability is generated from model structured-output metadata; custom Bedrock models can override compat.supportsStrictMode. OpenAI Responses and Chat Completions can also emit grammar-constrained custom tools with OpenAI Lark or regex grammar variants. If multiple OpenAI variants are supplied, Lark is preferred over regex. Grammar constraints are enforced when the active model supports grammar tools; otherwise the tool falls back to normal function/JSON-schema handling. Grammar tool capability is model metadata: the generated catalog sets compat.supportsOpenAIGrammarTools for GPT-5+ models on endpoints that pass OpenAI custom tools through (OpenAI, OpenAI Codex, Azure OpenAI Responses, GitHub Copilot, opencode, and Cloudflare AI Gateway). OpenAI rejects type: "custom" tools for pre-GPT-5 models, and gateways that normalize tool schemas (e.g. OpenRouter) mangle them, so the flag stays off elsewhere. Custom model definitions can opt in via compat. Grammar-capable models reject grammar configurations without a non-empty supported variant. Native grammar tools must have an object parameter schema with exactly one required string property:

const patchTool: Tool = {
  name: 'apply_patch',
  description: 'Apply a patch',
  parameters: Type.Object({
    input: Type.String()
  }, { additionalProperties: false }),
  constrainedSampling: {
    type: 'grammar',
    variants: {
      openai_lark: 'start: /.+/s'
    }
  }
};

Handling Tool Calls

Tool results use content blocks and can include both text and images:

import { readFileSync } from 'fs';

const context: Context = {
  messages: [{ role: 'user', content: 'What is the weather in London?', timestamp: Date.now() }],
  tools: [weatherTool]
};

const response = await models.complete(model, context);

// Check for tool calls in the response
for (const block of response.content) {
  if (block.type === 'toolCall') {
    // Execute your tool with the arguments
    // See "Validating Tool Arguments" section for validation
    const result = await executeWeatherApi(block.arguments);

    // Add tool result with text content
    context.messages.push({
      role: 'toolResult',
      toolCallId: block.id,
      toolName: block.name,
      content: [{ type: 'text', text: JSON.stringify(result) }],
      isError: false,
      timestamp: Date.now()
    });
  }
}

// Tool results can also include images (for vision-capable models)
const imageBuffer = readFileSync('chart.png');
context.messages.push({
  role: 'toolResult',
  toolCallId: 'tool_xyz',
  toolName: 'generate_chart',
  content: [
    { type: 'text', text: 'Generated chart showing temperature trends' },
    { type: 'image', data: imageBuffer.toString('base64'), mimeType: 'image/png' }
  ],
  isError: false,
  timestamp: Date.now()
});

Streaming Tool Calls with Partial JSON

During streaming, tool call arguments are progressively parsed as they arrive. This enables real-time UI updates before the complete arguments are available:

const s = models.stream(model, context);

for await (const event of s) {
  if (event.type === 'toolcall_delta') {
    const toolCall = event.partial.content[event.contentIndex];

    // toolCall.arguments contains partially parsed JSON during streaming
    // This allows for progressive UI updates
    if (toolCall.type === 'toolCall' && toolCall.arguments) {
      // BE DEFENSIVE: arguments may be incomplete
      // Example: Show file path being written even before content is complete
      if (toolCall.name === 'write_file' && toolCall.arguments.path) {
        console.log(`Writing to: ${toolCall.arguments.path}`);

        // Content might be partial or missing
        if (toolCall.arguments.content) {
          console.log(`Content preview: ${toolCall.arguments.content.substring(0, 100)}...`);
        }
      }
    }
  }

  if (event.type === 'toolcall_end') {
    // Here toolCall.arguments is complete (but not yet validated)
    const toolCall = event.toolCall;
    console.log(`Tool completed: ${toolCall.name}`, toolCall.arguments);
  }
}

Important notes about partial tool arguments:

  • During toolcall_delta events, arguments contains the best-effort parse of partial JSON
  • Fields may be missing or incomplete - always check for existence before use
  • String values may be truncated mid-word
  • Arrays may be incomplete
  • Nested objects may be partially populated
  • At minimum, arguments will be an empty object {}, never undefined
  • The Google provider does not support function call streaming. Instead, you will receive a single toolcall_delta event with the full arguments.

Validating Tool Arguments

When implementing your own tool execution loop, use validateToolCall to validate arguments before passing them to your tools:

import { validateToolCall, type Tool } from '@earendil-works/pi-ai';

const tools: Tool[] = [weatherTool, calculatorTool];
const s = models.stream(model, { messages, tools });

for await (const event of s) {
  if (event.type === 'toolcall_end') {
    const toolCall = event.toolCall;

    try {
      // Validate arguments against the tool's schema (throws on invalid args)
      const validatedArgs = validateToolCall(tools, toolCall);
      const result = await executeMyTool(toolCall.name, validatedArgs);
      // ... add tool result to context
    } catch (error) {
      // Validation failed - return error as tool result so model can retry
      context.messages.push({
        role: 'toolResult',
        toolCallId: toolCall.id,
        toolName: toolCall.name,
        content: [{ type: 'text', text: error.message }],
        isError: true,
        timestamp: Date.now()
      });
    }
  }
}

Complete Event Reference

Successful generation follows start → updates* → done. A failure after generation starts follows start → updates* → error. Request setup may fail before generation starts, in which case the stream contains only error; done and update events are invalid before start. Direct API streamSimple() calls throw synchronously when request auth is missing.

Every non-terminal event's partial is the shared live response-so-far helper. It is intentionally not an event-time snapshot: providers may mutate the same message and content blocks as generation advances, including while older events wait in the stream queue. Inspect it when handling an event instead of retaining it as historical state. Text and ordinary thinking blocks are empty when their *_start event is emitted and grow only through matching *_delta events until the authoritative *_end; redacted thinking may be complete at start and emit no deltas. Tool-call arguments at toolcall_start are provider-specific; toolcall_delta carries subsequent JSON updates.

All streaming events emitted during assistant message generation:

| Event Type | Description | Key Properties | |------------|-------------|----------------| | start | Stream begins | partial: Initial assistant message structure | | text_start | Text block starts | contentIndex: Position in content array | | text_delta | Text chunk received | delta: New text, contentIndex: Position | | text_end | Text block complete | content: Full text, contentIndex: Position | | thinking_start | Thinking block starts | contentIndex: Position in content array | | thinking_delta | Thinking chunk received | delta: New text, contentIndex: Position | | thinking_end | Thinking block complete | content: Full thinking, contentIndex: Position | | toolcall_start | Tool call begins | contentIndex: Position in content array | | toolcall_delta | Tool arguments streaming | delta: JSON chunk, partial.content[contentIndex].arguments: Partial parsed args | | toolcall_end | Tool call complete | toolCall: Complete, but not schema-validated, tool call with id, name, arguments | | done | Stream complete | reason: Stop reason ("stop", "length", "toolUse"), message: Final assistant message | | error | Error occurred | reason: Error type ("error" or "aborted"), error: AssistantMessage with partial content |

Streaming events for different content blocks are not guaranteed to be contiguous. Providers may emit deltas for text, thinking, and tool calls in the same upstream chunk, and pi may surface corresponding events interleaved, for example text_start, text_delta, toolcall_start, text_delta, toolcall_delta. Consumers must use contentIndex to associate each delta/end event with its block and must not assume that a block's *_start/*_delta/*_end sequence is uninterrupted by events for other blocks.

Compact Assistant Message Frames

AssistantMessageFrameEncoder converts one stream into compact, persistable AssistantMessageFrame values. Create one encoder per stream and feed it every event in order. The encoder understands that partial is live: a block-start event consumed after the provider has already queued later deltas snapshots the current block once, and covered queued text/thinking deltas produce no duplicate frame. It retains only per-open-block counters plus, temporarily, the raw prefix needed to synchronize an already-advanced tool call. It never clones the growing full partial per token.

The start frame contains message metadata with empty content. Text and thinking frames store each generated character at most once before the authoritative end frame. Tool calls that were already advanced when their start event was consumed use one compact JSON checkpoint before ordinary deltas resume. Terminal done and error events produce no frame because final message settlement is separate. A pre-generation error therefore produces no frames.

reduceAssistantMessageFrames() is the canonical pure reducer. It reconstructs text, thinking, and tool-call arguments, including interleaved blocks identified by contentIndex, and rejects malformed sequences. It performs a single pass over the iterable and returns undefined when there is no start frame. End frames replace blocks with the provider's authoritative completed content and metadata. The reducer does not validate tool arguments against a TypeBox schema; call validateToolCall before execution.

import {
  AssistantMessageFrameEncoder,
  reduceAssistantMessageFrames,
  type AssistantMessageFrame,
} from '@earendil-works/pi-ai';

const encoder = new AssistantMessageFrameEncoder();
const frames: AssistantMessageFrame[] = [];
for await (const event of s) {
  const frame = encoder.encode(event);
  if (frame) frames.push(frame);
}

const reconstructedPartial = reduceAssistantMessageFrames(frames);
const finalMessage = await s.result(); // Persist terminal settlement separately.

An encoder rejects duplicate starts, updates before start, done before start, events after a terminal event, duplicate block starts, and block-kind mismatches. An error before start is valid and returns no frame.

Image Input

Models with vision capabilities can process images. You can check if a model supports images via the input property. If you pass images to a non-vision model, they are silently ignored.

import { readFileSync } from 'fs';

const model = models.getModel('openai', 'gpt-4o-mini')!;

// Check if model supports images
if (model.input.includes('image')) {
  console.log('Model supports vision');
}

const imageBuffer = readFileSync('image.png');
const base64Image = imageBuffer.toString('base64');

const response = await models.complete(model, {
  messages: [{
    role: 'user',
    content: [
      { type: 'text', text: 'What is in this image?' },
      { type: 'image', data: base64Image, mimeType: 'image/png' }
    ],
    timestamp: Date.now()
  }]
});

// Access the response
for (const block of response.content) {
  if (block.type === 'text') {
    console.log(block.text);
  }
}

Image Generation

Image models live in the same Models collection and on the same Provider as chat models, so one credential per provider covers both. They are typed ImageModel with type: "image" and are used through generateImages(), a one-shot API that waits for the provider response and returns the final AssistantImages result. Do not use the chat/stream APIs for them; stream() rejects image models.

Basic Image Generation

import { builtinModels } from '@earendil-works/pi-ai/providers/all';

const models = builtinModels();

const model = models.getModelOfType('image', 'openrouter', 'google/gemini-2.5-flash-image')!;

// Auth resolves through the provider (OPENROUTER_API_KEY here); explicit apiKey wins
const result = await models.generateImages(model, {
  input: [{ type: 'text', text: 'Generate a red circle on a plain white background.' }]
});

for (const block of result.output) {
  if (block.type === 'text') {
    console.log(block.text);
  } else if (block.type === 'image') {
    console.log(block.mimeType);
    console.log(block.data.substring(0, 32));
  }
}

generateImages() accepts only ImageModel values. If an upstream model supports both chat and image generation, the catalog contains separate entries with the same provider and ID: getModel() returns its chat operation and getModelOfType('image', ...) returns its image operation. Failures never reject; they return an AssistantImages with stopReason: "error", including unknown providers, unconfigured auth, and providers without an image implementation.

A provider declares image support with the images option of createProvider(): a map from model.api to an implementation with generateImages(). Image models go into the same models list as chat models. api becomes optional when images is present, so an image-only provider is just a provider without chat models:

import { createProvider, envApiKeyAuth } from '@earendil-works/pi-ai';

const pixels = createProvider({
  id: 'pixels',
  auth: { apiKey: envApiKeyAuth('Pixels API key', ['PIXELS_API_KEY']) },
  models: [{
    type: 'image',
    id: 'flux-pro',
    name: 'FLUX Pro',
    api: 'pixels-images',
    provider: 'pixels',
    baseUrl: 'https://api.pixels.test/v1',
    input: ['text'],
    output: ['image'],
    cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
  }],
  images: {
    'pixels-images': { generateImages: async (model, context, options) => { /* ... */ } },
  },
});
models.setProvider(pixels);

The old global API (getImageModel() / getImageModels() / getImageProviders() / generateImages()) remains available on the compat entrypoint:

import { getImageModel, generateImages } from '@earendil-works/pi-ai/compat';

const model = getImageModel('openrouter', 'google/gemini-2.5-flash-image');
const result = await generateImages(model, {
  input: [{ type: 'text', text: 'Generate a red circle on a plain white background.' }]
}, {
  apiKey: process.env.OPENROUTER_API_KEY
});

Some models also support image input:

import { readFileSync } from 'fs';

const imageBuffer = readFileSync('input.png');
const result = await models.generateImages(model, {
  input: [
    { type: 'text', text: 'Create a variation of this image with a blue background.' },
    { type: 'image', data: imageBuffer.toString('base64'), mimeType: 'image/png' }
  ]
});

Check capabilities on the model metadata:

console.log(model.input);  // ['text'] or ['text', 'image']
console.log(model.output); // ['image'] or ['image', 'text']

Notes and Limitations

  • Image models and chat models share Models and Provider; list them with getModelsOfType('image') and run them with generateImages(), never the chat/stream APIs.
  • Image-generation models do not participate in tool calling.
  • Outputs are returned in AssistantImages.output and can include both base64-encoded ImageContent blocks and TextContent blocks.
  • Some models return only images, others return images plus text. Check model.output.
  • Some models accept image input, others are text-to-image only. Check model.input.
  • Like the streaming APIs, image generation supports options such as apiKey, signal, headers, onPayload, and onResponse, and results may include stopReason, responseId, and usage.
  • If you want a model to analyze images in a conversation or call tools, use the regular chat APIs with a model that supports image input.
  • At the moment, image generation is available through only one provider, OpenRouter.

Classification

Classifier models consume structured JSON state and answer one or more typed questions. They do not use chat or image-generation APIs. TypeSafe's Jev model is available from these built-in providers:

| Provider | Model IDs | Auth | | --- | --- | --- | | typesafe | jev-latest | TYPESAFE_API_KEY | | openrouter | typesafe/jev-1.13, ~typesafe/jev-latest | OPENROUTER_API_KEY or OpenRouter OAuth | | cloudflare-workers-ai | typesafe/jev | CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID | | vercel-ai-gateway | typesafe-ai/jev | AI_GATEWAY_API_KEY | | opencode | jev-1.13, jev-1.13-free | OPENCODE_API_KEY |

import { builtinModels } from '@earendil-works/pi-ai/providers/all';

const models = builtinModels();
const model = models.getModelOfType('classifier', 'typesafe', 'jev-latest')!;
const result = await models.classify(model, {
  state: { message: 'The change works perfectly, thanks.' },
  questions: {
    category: {
      type: 'choice',
      instructions: 'Classify the message.',
      criteria: {
        approval: 'The user approves of the result',
        correction: 'The user requests a correction'
      }
    },
    satisfaction: {
      type: 'score',
      instructions: 'Score user satisfaction.',
      criteria: ['dissatisfied', 'neutral', 'satisfied']
    },
    approved: {
      type: 'bool',
      instructions: 'Does the user approve?',
      criteria: { true: 'Approval', false: 'No approval' }
    }
  }
});

console.log(result.answers);

The public contract uses bool questions and { type: "bool", probability } answers. The TypeSafe adapter translates those to and from its noul wire representation. Like image generation, classify() resolves to a result with stopReason: "error" instead of rejecting for provider, authentication, or response errors.

When the service reports token counts, result.usage carries them with their cost at the model's catalog price, the same Usage shape as chat messages. All System One services report token counts; a request that was answered with malformed answers keeps its usage. Local classifiers such as llama-cpp-classify report no usage.

ClassifierOptions.temperature divides the answer logits by the given value before they are normalized; values above 1 soften the distribution. APIs that cannot apply it, such as System One, ignore it.

Chat models on llama.cpp

The llama-cpp-classify API turns a chat model served by llama.cpp's llama-server into a classifier. Each question becomes one chat prompt: the state, every question of the request, the state again, and the question with its answers under single-token labels (letters for a choice, Yes/No for a bool, digits for a score). The prompt up to the final question is shared by all questions of a request, so the server's prompt cache evaluates the state once per request. The server returns the log-probabilities of the next token, and the answer is the softmax over the label tokens. Choices support up to 62 options and scores up to 10 levels. The model's baseUrl is the server URL; a trailing /v1 is ignored. In router mode, the model ID selects the model.

import { createProvider } from '@earendil-works/pi-ai';
import { llamaCppClassifyApi } from '@earendil-works/pi-ai/api/llama-cpp-classify.lazy';

const provider = createProvider({
  id: 'local-llama',
  auth: { apiKey: { name: 'llama.cpp', resolve: async () => ({ auth: {} }) } },
  models: [{
    type: 'classifier',
    id: 'qwen3-4b',
    name: 'Qwen3 4B',
    api: 'llama-cpp-classify',
    provider: 'local-llama',
    baseUrl: 'http://127.0.0.1:8080',
    input: ['text'],
    cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
    contextWindow: 32768
  }],
  classifiers: { 'llama-cpp-classify': llamaCppClassifyApi() }
});

Raw label probabilities are usually overconfident; pass temperature above 1 to soften them.

Custom providers register classifier models and implementations by API ID:

createProvider({
  id: 'classifier-service',
  auth,
  models: [model],
  classifiers: {
    'classifier-api': { classify: async (model, context, options) => result }
  }
});

Thinking/Reasoning

Many models support thinking/reasoning capabilities where they can show their internal thought process. You can check if a model supports reasoning via the reasoning property. If you pass reasoning options to a non-reasoning model, they are silently ignored.

Unified Interface (streamSimple/completeSimple)

// Many models across providers support thinking/reasoning
const model = models.getModel('anthropic', 'claude-sonnet-4-5')!;
// or models.getModel('openai', 'gpt-5-mini');
// or models.getModel('google', 'gemini-2.5-flash');
// or models.getModel('xai', 'grok-4.7');

// Check if model supports reasoning
if (model.reasoning) {
  console.log('Model supports reasoning/thinking');
}

// Use the simplified reasoning option
const response = await models.completeSimple(model, {
  messages: [{ role: 'user', content: 'Solve: 2x + 5 = 13', timestamp: Date.now() }]
}, {
  reasoning: 'medium'  // 'minimal' | 'low' | 'medium' | 'high' | 'xhigh' | 'max'
});

// Access thinking and text blocks
for (const block of response.content) {
  if (block.type === 'thinking') {
    console.log('Thinking:', block.thinking);
  } else if (block.type === 'text') {
    console.log('Response:', block.text);
  }
}

xhigh and max are model-specific, opt-in levels. Use getSupportedThinkingLevels(model) to determine whether a concrete model exposes either level; models such as GPT-5.6 can expose both.

Provider-Specific Options (stream/complete)

models.stream()/complete() accept the owning API's full option set. Use hasApi() to narrow a dynamically looked-up model to its API for full option typing:

import { hasApi } from '@earendil-works/pi-ai';

// OpenAI Reasoning (o1, o3, gpt-5)
const openaiModel = models.getModel('openai', 'gpt-5-mini')!;
if (hasApi(openaiModel, 'openai-responses')) {
  await models.complete(openaiModel, context, {
    reasoningEffort: 'medium',
    reasoningSummary: 'detailed'  // OpenAI Responses API only
  });
}

// Anthropic Thinking
const anthropicModel = models.getModel('anthropic', 'claude-sonnet-4-5')!;
if (hasApi(anthropicModel, 'anthropic-messages')) {
  await models.complete(anthropicModel, context, {
    thinkingEnabled: true,
    thinkingBudgetTokens: 8192  // Optional token limit
  });
}

// Google Gemini Thinking
const googleModel = models.getModel('google', 'gemini-2.5-flash')!;
if (hasApi(googleModel, 'google-generative-ai')) {
  await models.complete(googleModel, context, {
    thinking: {
      enabled: true,
      budgetTokens: 8192  // -1 for dynamic, 0 to disable
    }
  });
}

Streaming Thinking Content

When streaming, thinking content is delivered through specific events:

const s = models.streamSimple(model, context, { reasoning: 'high' });

for await (const event of s) {
  switch (event.type) {
    case 'thinking_start':
      console.log('[Model started thinking]');
      break;
    case 'thinking_delta':
      process.stdout.write(event.delta);  // Stream thinking content
      break;
    case 'thinking_end':
      console.log('\n[Thinking complete]');
      break;
  }
}

Stop Reasons

Every AssistantMessage includes a stopReason field that indicates how the generation ended:

  • "pending" - Only present in partial messages when we do not know what the stop reason will be
  • "stop" - This is the final message the model will produce this turn
  • "length" - Output hit the maximum token limit
  • "toolUse" - Model is calling tools and expects tool results
  • "error" - An error occurred during generation
  • "aborted" - Request was cancelled via abort signal

AssistantMessage may also include responseId, a provider-specific upstream response or message identifier when the underlying API exposes one. Do not assume it is always present across providers.

Error Handling

Request failures after a stream is returned never throw: when a request ends with an error (including aborts and tool call validation errors), the streaming API emits an error event and the final message carries the details. Setup failures may emit error without start; failures after generation begins emit start, any observed updates, then error. Direct API streamSimple() calls throw synchronously when request auth is missing:

// In streaming
for await (const event of s) {
  if (event.type === 'error') {
    // event.reason is either "error" or "aborted"
    // event.error is the AssistantMessage with partial content
    console.error(`Error (${event.reason}):`, event.error.errorMessage);
    console.log('Partial content:', event.error.content);
  }
}

// The final message will have the error details
const message = await s.result();
if (message.stopReason === 'error' || message.stopReason === 'aborted') {
  console.error('Request failed:', message.errorMessage);
  // message.content contains any partial content received before the error
  // message.usage contains partial token counts and costs
}

When using a provider collection, auth failures (OAuth refresh failed, unknown provider) surface as a stream error with stopReason: "error". Direct API streamSimple() calls instead throw synchronously when their required auth is absent.

Aborting Requests

The abort signal allows you to cancel in-progress requests. Aborted requests have stopReason === 'aborted':

const controller = new AbortController();

// Abort after 2 seconds
setTimeout(() => controller.abort(), 2000);

const s = models.stream(model, {
  messages: [{ role: 'user', content: 'Write a long story', timestamp: Date.now() }]
}, {
  signal: controller.signal
});

for await (const event of s) {
  if (event.type === 'text_delta') {
    process.stdout.write(event.delta);
  } else if (event.type === 'error') {
    // event.reason tells you if it was "error" or "aborted"
    console.log(`${event.reason === 'aborted' ? 'Aborted' : 'Error'}:`, event.error.errorMessage);
  }
}

// Get results (may be partial if aborted)
const response = await s.result();
if (response.stopReason === 'aborted') {
  console.log('Request was aborted:', response.errorMessage);
  console.log('Partial content received:', response.content);
  console.log('Tokens used:', response.usage);
}

Continuing After Abort

Aborted messages can be added to the conversation context and continued in subsequent requests:

const context = {
  messages: [
    { role: 'user', content: 'Explain quantum computing in detail', timestamp: Date.now() }
  ]
};

// First request gets aborted after 2 seconds
const controller1 = new AbortController();
setTimeout(() => controller1.abort(), 2000);

const partial = await models.complete(model, context, { signal: controller1.signal });

// Add the partial response to context
context.messages.push(partial);
context.messages.push({ role: 'user', content: 'Please continue', timestamp: Date.now() });

// Continue the conversation
const continuation = await models.complete(model, context);

Debugging Provider Payloads

Use the onPayload callback to inspect the request payload sent to the provider. This is useful for debugging request formatting issues or provider validation errors.

const response = await models.complete(model, context, {
  onPayload: (payload) => {
    console.log('Provider payload:', JSON.stringify(payload, null, 2));
  }
});

The callback is supported by stream, complete, streamSimple, and completeSimple.

Observing Provider Stream Events

Use onProviderStreamEvent to inspect provider-specific fields that Pi does not include in AssistantMessage. The callback receives the parsed event available to the adapter before Pi normalizes it. Treat the event as read-only because mutations can affect normalization. This is not guaranteed to be the original HTTP bytes or SSE frame.

const openRouterModel = models.getModel('openrouter', 'openrouter/auto')!;
const response = await models.complete(openRouterModel, context, {
  headers: { "X-OpenRouter-Metadata": "enabled" },
  onProviderStreamEvent: (data) => {
    const chunk = data as Record<string, unknown>;
    if (chunk.openrouter_metadata) {
      console.log(chunk.openrouter_metadata);
    }
  },
});

Callbacks are awaited in stream order, so slow callbacks delay stream consumption and thrown errors fail the request. SDK-backed adapters can expose only fields retained by their SDK.

Custom Providers

createProvider()

createProvider() builds a provider from parts: identity, auth, a model list, and an API implementation (api for chat models, images for image generation, classifiers for classification; at least one is required, see Image Generation). Use it for local inference servers, proxies, or any OpenAI/Anthropic-compatible endpoint:

import { createModels, createProvider, envApiKeyAuth, type Model } from '@earendil-works/pi-ai';
import { openAICompletionsApi } from '@earendil-works/pi-ai/api/openai-completions.lazy';

const ollamaModel: Model<'openai-completions'> = {
  id: 'llama-3.1-8b',
  name: 'Llama 3.1 8B (Ollama)',
  api: 'openai-completions',
  provider: 'ollama',
  baseUrl: 'http://localhost:11434/v1',
  reasoning: false,
  input: ['text'],
  cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
  contextWindow: 128000,
  maxTokens: 32000
};

const ollama = createProvider({
  id: 'ollama',
  name: 'Ollama',
  baseUrl: 'http://localhost:11434/v1',
  // Every provider declares auth; keyless local servers resolve as configured with no key.
  auth: { apiKey: { name: 'Ollama', resolve: async () => ({ auth: {} }) } },
  models: [ollamaModel],
  api: openAICompletionsApi(),
});

const models = createModels();
models.setProvider(ollama);

await models.complete(models.getModel('ollama', 'llama-3.1-8b')!, context);

For providers with real keys, envApiKeyAuth(displayName, envVars) gives the standard behavior (stored credential wins, then the first set env var):

const proxy = createProvider({
  id: 'my-proxy',
  auth: { apiKey: envApiKeyAuth('My proxy API key', ['MY_PROXY_API_KEY']) },
  models: [/* ... */],
  api: openAICompletionsApi(),
});

Mixed-API providers pass a map keyed by model.api; each model dispatches to its API's implementation:

import { anthropicMessagesApi } from '@earendil-works/pi-ai/api/anthropic-messages.lazy';
import { openAIResponsesApi } from '@earendil-works/pi-ai/api/openai-responses.lazy';

const gateway = createProvider({
  id: 'my-gateway',
  auth: { apiKey: envApiKeyAuth('Gateway key', ['GATEWAY_API_KEY']) },
  models: [/* models with api: 'anthropic-messages' or 'openai-responses' */],
  api: {
    'anthropic-messages': anthropicMessagesApi(),
    'openai-responses': openAIResponsesApi(),
  },
});

Provider-wide endpoint or request transformations belong in the provider's API implementation: wrap the ProviderStreams you pass as api so every request goes through the transformation before dispatch. The Cloudflare providers do this to materialize account/gateway endpoint placeholders from the resolved provider env:

function tenantStreams(streams: ProviderStreams): ProviderStreams {
  const withTenant = (model: Model<Api>) => ({ ...model, baseUrl: model.baseUrl.replace('{tenant}', tenantId) });
  return {
    stream: (model, context, options) => streams.stream(withTenant(model), context, options),
    streamSimple: (model, context, options) => streams.streamSimple(withTenant(model), context, options),
  };
}

const tenantGateway = createProvider({
  id: 'tenant-gateway',
  auth: { apiKey: envApiKeyAuth('Gateway key', ['GATEWAY_API_KEY']) },
  models: [/* ... */],
  api: tenantStreams(openAICompletionsApi()),
});

Dynamic model lists use fetchModels, which can return models of every type. Models.refresh() refreshes every configured dynamic provider, passing its effective API-key or refreshed OAuth credential. A ModelsStore persists dynamic catalogs; both stores default to in-memory implementations. Its read, write, and delete operations accept optional cancellation, and Models binds those waits to the provider refresh signal.

const models = createModels({ credentials, modelsStore });
const llamacpp = createProvider({
  id: 'llamacpp',
  auth: { apiKey: { name: 'llama.cpp', resolve: async () => ({ auth: {} }) } },
  models: [],
  fetchModels: async ({ signal }) => fetchModelsFromServer('http://localhost:8080', signal),
  api: openAICompletionsApi(),
});

models.setProvider(llamacpp);
const result = await models.refresh({ signal });
if (result.aborted) console.log('refresh cancelled');
for (const [provider, error] of result.errors) console.error(provider, error);

Models.refresh() is unbounded when its optional signal is omitted. Providers always receive a concrete RefreshModelsContext.signal and must honor it for network requests and other blocking work. When a caller supplies a signal, Models.refresh() returns promptly with aborted: true after cancellation even if a custom provider fails to cooperate; the provider must still honor the signal to stop its underlying work.

Use models.refresh({ providers: ['openrouter'] }) to restrict work to selected providers, models.refresh({ allowNetwork: false }) to restore persisted catalogs without network access, or models.refresh({ force: true }) to bypass provider freshness checks. Model reads stay synchronous and return the last restored or refreshed list.

createProvider() handles dynamic publication and persistence automatically. Handwritten Provider.refreshModels() implementations receive the read-only context.stored snapshot and publish through context.publish({ persist?, update? }). Omit persist to leave storage unchanged, pass a ModelsStoreEntry to write it, or pass persist: null to delete it. ModelsStoreEntry.models contains models of every type. Publication is generation-checked; put synchronous in-memory catalog changes in update rather than mutating state before publication.

Custom models can carry headers (e.g. proxies behind bot detection) and compat flags. Models.getAuth(model) includes those model headers, and stream methods merge them before explicit request headers and transformHeaders. See OpenAI Compatibility Settings.

Some OpenAI-compatible servers do not understand the developer role used for reasoning-capable models. For those providers, set compat.supportsDeveloperRole to false so the system prompt is sent as a system message instead. If the server also does not support reasoning_effort, set compat.supportsReasoningEffort to false too. This commonly applies to Ollama, vLLM, SGLang, and similar OpenAI-compatible servers.

Use model-level thinkingLevelMap to describe model-specific thinking controls. Keys are pi thinking levels (off, minimal, low, medium, high, xhigh, max). Missing standard levels through high use provider defaults; xhigh and max are opt-in and require a non-null map entry. String values are sent to the provider, null marks a level unsupported, and maps may skip levels.

const ollamaReasoningModel: Model<'openai-completions'> = {
  id: 'gpt-oss:20b',
  name: 'GPT-OSS 20B (Ollama)',
  api: 'openai-completions',
  provider: 'ollama',
  baseUrl: 'http://localhost:11434/v1',
  reasoning: true,
  input: ['text'],
  cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
  contextWindow: 131072,
  maxTokens: 32000,
  thinkingLevelMap: {
    minimal: null,
    low: null,
    medium: null,
    high: 'high',
    xhigh: null,
  },
  compat: {
    supportsDeveloperRole: false,
    supportsReasoningEffort: false,
  }
};

Calling API Implementations Directly

The API implementations are importable on their own. Each module exports exactly stream and streamSimple with that API's full option typing. Direct calls bypass provider auth and context normalization — pass apiKey explicitly and wrap the context in normalizeContext():

import { normalizeContext } from '@earendil-works/pi-ai';
import { stream } from '@earendil-works/pi-ai/api/anthropic-messages';

const s = stream(claudeModel, normalizeContext(context), {
  apiKey: process.env.ANTHROPIC_API_KEY,
  thinkingEnabled: true,
  thinkingBudgetTokens: 2048,
});

Built-in API implementations live under ./api/<api-id>:

| API id | Options type | |--------|--------------| | anthropic-messages | AnthropicOptions | | openai-completions | OpenAICompletionsOptions | | openai-responses | OpenAIResponsesOptions | | openai-codex-responses | OpenAICodexResponsesOptions | | azure-openai-responses | AzureOpenAIResponsesOptions | | google-generative-ai | GoogleOptions | | google-vertex | GoogleVertexOptions | | mistral-conversations | MistralOptions | | bedrock-converse-stream | BedrockOptions |

Importing an implementation module loads its SDK. The ./api/<id>.lazy wrappers (used by the provider factories) defer that load to the first request when the runtime or bundler supports dynamic import chunking. Legacy raw API subpaths from older releases (./anthropic, ./google, ./mistral, ./openai-completions, ...) were removed; use @earendil-works/pi-ai/api/<api-id>.

OpenAI Compatibility Settings

The openai-completions API is implemented by many providers with minor differences. By default, the library auto-detects compatibility settings based on baseUrl for a small set of known OpenAI-compatible providers (Cerebras, xAI, Chutes, DeepSeek, NVIDIA NIM, Together AI, zAi, OpenCode, Cloudflare Workers AI, etc.). For custom proxies or unknown endpoints, you can override these settings via the compat field. For openai-responses models, the compat field supports Responses-specific flags.

interface OpenAICompletionsCompat {
  supportsStore?: boolean;           // Whether provider supports the `store` field (default: true)
  supportsDeveloperRole?: boolean;   // Whether provider supports `developer` role vs `system` (default: true)
  supportsReasoningEffort?: boolean; // Whether provider supports `reasoning_effort` (default: true)
  supportsUsageInStreaming?: boolean; // Whether provider supports `stream_options: { include_usage: true }` (default: true)
  supportsStrictMode?: boolean;      // Whether provider supports `strict` in tool definitions (default: false; enabled in metadata for capable built-in models)
  supportsOpenAIGrammarTools?: boolean; // Whether to emit OpenAI custom Lark/regex grammar tools; false falls back to normal function tools (default: false; the generated catalog enables it for capable models)
  supportsMidConvoSystemMessages?: boolean; // Whether the model accepts system messages after the conversation started; false folds them into