npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@aituber-onair/core

v0.26.18

Published

Core library for AITuber OnAir providing voice synthesis and chat processing

Readme

AITuber OnAir Core

AITuber OnAir Core - logo

AITuber OnAir Core is a TypeScript library developed to provide functionality for the AITuber OnAir web service, designed for AI-based virtual streaming (AITuber).

日本語版はこちら

While it is primarily intended to provide functionality for AITuber OnAir, this project is published as open-source software and is available as an npm package under the MIT License.

It specializes in generating response text and audio from text or image inputs, and is designed to easily integrate with other parts of an application (storage, YouTube integration, avatar control, etc.).

Chat and Voice model updates

Core exposes the models and capability helpers from Chat 0.61.0 and the speech engines, option types, and endpoint helpers from Voice 0.25.0. Existing provider defaults are unchanged.

  • Native models: GPT-6 Astra, Sol, and Luna; Claude Fable 5.1, Opus 5.5, and Sonnet 5.5; DeepSeek V4.1 Flash; and Grok 4.7.
  • OpenAI-compatible endpoint helpers: resolveOpenAICompatibleEndpoint, listOpenAICompatibleModels, testOpenAICompatibleConnection, and OPENAI_COMPATIBLE_LOCAL_PRESETS (Ollama, LM Studio, llama.cpp, vLLM). See the local LLM guide.
  • OpenAI-compatible speech helpers: resolveOpenAICompatibleSpeechEndpoint, listOpenAICompatibleSpeechModels, listOpenAICompatibleSpeechVoices, getOpenAICompatibleSpeechServerInfo, and testOpenAICompatibleSpeech for local TTS servers. See the local TTS guide.
  • OpenRouter: GPT-6 Astra / Astra Pro, Claude Fable 5.1, DeepSeek V4.1 Flash, Gemini 3.8 Flash, Ling 3.0 Flash VL (free), Mercury 2.5, Nex N2.5 Mini / Pro (free), Qwen3.8 Max 0902, and Muse Spark 1.3.
  • Models with required reasoning use their supported effort values. Native Astra uses the Responses API and starts at low; use isOpenAIReasoningModel and getDefaultReasoningEffortForOpenAIModel for the expanded OpenAI family. Existing GPT-5 helpers remain available.
  • Native DeepSeek tool calling requires non-thinking mode (reasoning_effort: 'none'). OpenRouter reasoning and vision support follow the selected model's Chat capability metadata.
  • Gemini TTS adds gemini-3.8-flash-lite-tts and gemini-3.8-flash-tts to every React example. Existing TTS defaults remain unchanged. On 3.8 models, the style prompt is sent as speech metadata and language code is unavailable.
  • Inworld supports inworld-tts-2-flash in addition to the default inworld-tts-2. Models without delivery-mode support, currently Flash, omit that option and disable its control in all React examples. Voice-list selection, persisted settings, and other audio options remain available.

The React basic example lists the new models explicitly. Avatar examples load the supported model list from Core and apply the Casual preset to supported OpenAI reasoning models. Node examples pass the model and provider options through to Core.

  • New explicit Chat options include native GPT-6.1 Sol and Z.ai GLM-5.3 FlashX, Mistral-hosted zai-glm-5-3, and OpenRouter GPT-6.1 Sol / GPT-6 Sol / Luna, Claude Sonnet / Opus 5.5, Grok 4.7, GLM-5.3 FlashX, nvidia/nemotron-3.5-lightning, qwen/qwen3.8-27b, and qwen/qwen3.8-omni-flash.
  • Models without documented reasoning effort or token-budget controls omit those fields. OPENROUTER_MODELS_WITHOUT_REASONING_BUDGET identifies the three new Nemotron/Qwen options; the two Qwen models support images, while Nemotron and Mistral GLM-5.3 are text-only. Mistral GLM-5.3 requires an eligible subscription tier and has no adjustable reasoning effort.
  • GPT-6.1 Sol uses low reasoning for the casual preset. Responses is the default and is required for tools; tool-free Chat Completions supports low, medium, high, xhigh, and max. Z.ai GLM-5.3 FlashX always reasons, supports images, and requires a Model API key. Native Z.ai browser calls may fail because of provider CORS restrictions.
  • Voice options include ElevenLabs eleven_v4 (Stability and Similarity; no Style, Speed, or Speaker Boost), Cartesia sonic-3.6 and sonic-3.6-2026-08-27 (default remains sonic-3.5), Deepgram Flux, and the opt-in Gradium gradium-tts-beta model through gradiumModel.

For endpoint details and limitations, see the Chat README and Voice README.

Deepgram Flux setup

Use engineType: 'deepgram', an API key, and an English Flux voice ID such as flux-haley-en. Core re-exports DeepgramEngine, its option types, and the speech/catalog URL constants. getVoiceEngineVoiceList('deepgram') lists English voices from the public v2 catalog without an API key. deepgramSpeed accepts 0.5–1.5 in 0.05 steps. This integration returns a complete MP3 from POST /v2/speak; it does not use Aura or WebSocket speech. See the Deepgram batch guide.

All ten React TTS examples use /api/deepgram/v2/speak and /api/deepgram/v2/models via Vite dev/preview proxies. Production apps need equivalent authenticated backend routes with API keys kept server-side. Avatar settings allow custom TTS/catalog URLs. Voice-list refreshes preserve the selected voice and the previous list on empty or failed responses; Haley remains selectable. Gradium settings default to production and reset to production when switching back to Gradium.

Table of Contents

Overview

AITuberOnAirCore is the central module that provides core features for AI tubers. It forms the core of the AITuber OnAir application. It encapsulates complex AI response generation, conversation context management, speech synthesis, and more, making these features available through a simple API.

Installation

You can install AITuber OnAir Core using npm:

npm install @aituber-onair/core

Or using yarn:

yarn add @aituber-onair/core

Or using pnpm:

pnpm install @aituber-onair/core

Main Features

  • AI Response Generation from Text Input
    Generates natural responses to user text input using OpenAI GPT models.
  • AI Response Generation from Images (Vision)
    Generates AI responses based on recognized content from images (e.g., live broadcast screens).
  • Conversation Context Management & Memory
    Maintains long-running conversation context via short-, mid-, and long-term memory systems.
  • Text-to-Speech Conversion
    Compatible with multiple speech engines (VOICEVOX, VoicePeak, AivisSpeech, Aivis Cloud, OpenAI TTS, Gemini TTS, xAI, Unreal Speech, ElevenLabs, Inworld, Gradium, Piper Plus, Web Speech API).
  • Emotion Extraction & Processing
    Extracts emotion from AI responses and utilizes it for speech synthesis or avatar expressions.
  • Event-Driven Architecture
    Emits events at each stage of processing to simplify external integrations.
  • Customizable Prompts
    Allows customization of prompts for vision processing and conversation summarization.
  • Pluggable Persistence
    Memory features can be persisted via LocalStorage, IndexedDB, or other customizable methods.
  • Function Calling with Tools Support
    Enables AI to use tools for performing actions beyond text generation, such as calculations, API calls, or data retrieval.

Basic Usage

Below is a simplified example of how to use AITuber OnAir Core:

import {
  AITuberOnAirCore,
  AITuberOnAirCoreEvent,
  AITuberOnAirCoreOptions
} from '@aituber-onair/core';

// 1. Define options
const options: AITuberOnAirCoreOptions = {
  chatProvider: 'openai', // Optional. If omitted, the default OpenAI will be used.
  apiKey: 'YOUR_API_KEY',
  chatOptions: {
    systemPrompt: 'You are an AI streamer. Act as a cheerful and friendly live broadcaster.',
    visionSystemPrompt: 'Please comment like a streamer on what is shown on screen.',
    visionPrompt: 'Look at the broadcast screen and provide commentary suited to the situation.',
    memoryNote: 'This is a summary of past conversations. Please refer to it appropriately to continue the conversation.',
    // Response length control
    maxTokens: 150,                    // Direct token limit for text chat
    responseLength: 'medium',          // Or use preset: 'veryShort', 'short', 'medium', 'long'
    visionMaxTokens: 200,              // Direct token limit for vision processing
    visionResponseLength: 'long',      // Or use preset for vision responses
  },
  // OpenAI Default model is gpt-4o-mini
  // You can specify different models for text chat and vision processing
  // model: 'o3-mini',        // Lightweight model for text chat (no vision support)
  // visionModel: 'gpt-4o',   // Model capable of image processing
  memoryOptions: {
    enableSummarization: true,
    shortTermDuration: 60 * 1000, // 1 minute
    midTermDuration: 4 * 60 * 1000, // 4 minutes
    longTermDuration: 9 * 60 * 1000, // 9 minutes
    maxMessagesBeforeSummarization: 20,
    maxSummaryLength: 256,
    // You can specify a custom summarization prompt
    summaryPromptTemplate: 'Please summarize the following conversation in under {maxLength} characters. Include important points.'
  },
  voiceOptions: {
    engineType: 'voicevox', // Speech engine type
    speaker: '1',           // Speaker ID
    apiKey: 'ENGINE_SPECIFIC_API_KEY', // If required (e.g., OpenAI, MiniMax)
    groupId: 'YOUR_GROUP_ID',          // If using MiniMax
    endpoint: 'global',                // If using MiniMax: 'global' or 'china'
    onComplete: () => console.log('Voice playback completed'),
    // Custom API endpoint URLs (optional)
    voicevoxApiUrl: 'http://custom-voicevox-server:50021',
    voicepeakApiUrl: 'http://custom-voicepeak-server:20202',
    aivisSpeechApiUrl: 'http://custom-aivis-server:10101',
  },
  debug: true, // Enable debug output
};

// 2. Create an instance
const aituber = new AITuberOnAirCore(options);

// 3. Set up event listeners
aituber.on(AITuberOnAirCoreEvent.PROCESSING_START, () => {
  console.log('Processing started');
});

aituber.on(AITuberOnAirCoreEvent.ASSISTANT_PARTIAL, (text) => {
  // Receive streaming responses and display in UI
  console.log(`Partial response: ${text}`);
});

aituber.on(AITuberOnAirCoreEvent.ASSISTANT_RESPONSE, (data) => {
  const { message, screenplay, rawText } = data;
  console.log(`Complete response: ${message.content}`);
  console.log(`Original text with emotion tags: ${rawText}`);
  if (screenplay.emotion) {
    console.log(`Emotion: ${screenplay.emotion}`);
  }
});

aituber.on(AITuberOnAirCoreEvent.SPEECH_START, (data) => {
  // The SPEECH_START event includes the screenplay object and rawText
  if (data && data.screenplay) {
    console.log(`Speech playback started: emotion = ${data.screenplay.emotion || 'neutral'}`);
    console.log(`Original text with emotion tags: ${data.rawText}`);
  } else {
    console.log('Speech playback started');
  }
});

aituber.on(AITuberOnAirCoreEvent.SPEECH_END, () => {
  console.log('Speech playback finished');
});

aituber.on(AITuberOnAirCoreEvent.TOOL_USE, (toolBlock) => 
  console.log(`Tool use -> ${toolBlock.name}`, toolBlock.input));

aituber.on(AITuberOnAirCoreEvent.TOOL_RESULT, (resultBlock) => 
  console.log(`Tool result ->`, resultBlock.content));

aituber.on(AITuberOnAirCoreEvent.ERROR, (error) => {
  console.error('Error occurred:', error);
});

// Memory and chat history related events
aituber.on(AITuberOnAirCoreEvent.CHAT_HISTORY_SET, (messages) => 
  console.log('Chat history set:', messages.length));

aituber.on(AITuberOnAirCoreEvent.CHAT_HISTORY_CLEARED, () => 
  console.log('Chat history cleared'));

aituber.on(AITuberOnAirCoreEvent.MEMORY_CREATED, (memory) => 
  console.log(`New memory created: ${memory.type}`));

aituber.on(AITuberOnAirCoreEvent.MEMORY_REMOVED, (memoryIds) => 
  console.log('Memory removed:', memoryIds));

aituber.on(AITuberOnAirCoreEvent.MEMORY_LOADED, (memories) => 
  console.log('Memory loaded:', memories.length));

aituber.on(AITuberOnAirCoreEvent.MEMORY_SAVED, (memories) => 
  console.log('Memory saved:', memories.length));

// 4. Process text input
await aituber.processChat('Hello, how is the weather today?');

// 5. Clear event listeners if needed
aituber.offAll();

Tool System

AITuber OnAir Core includes a powerful tool system that allows AI to perform actions beyond text generation, such as retrieving data or making calculations. This is particularly useful for creating interactive AITuber experiences.

Tool Definition Structure

Tools are defined using the ToolDefinition interface, which conforms to the function calling specification used by LLM providers:

type ToolDefinition = {
  name: string;                 // The name of the tool
  description?: string;         // Optional description of what the tool does
  parameters: {
    type: 'object';             // Must be 'object' (strictly typed)
    properties?: Record<string, {
      type?: string;            // Parameter type (e.g. 'string', 'integer')
      description?: string;     // Parameter description
      enum?: any[];             // For enumerated values
      items?: any;              // For array types
      required?: string[];      // Required nested properties
      [key: string]: any;       // Other JSON Schema properties
    }>;
    required?: string[];        // Names of required parameters
    [key: string]: any;         // Other JSON Schema properties
  };
  config?: { timeoutMs?: number }; // Optional configuration
};

Note that the parameters.type property is strictly typed as 'object' to conform to function calling standards used by LLM providers.

Registering and Using Tools

Tools are registered when initializing AITuberOnAirCore:

// Define a tool
const randomIntTool: ToolDefinition = {
  name: 'randomInt',
  description: 'Return a random integer from 0 to (max - 1)',
  parameters: {
    type: 'object',  // This must be 'object'
    properties: {
      max: {
        type: 'integer',
        description: 'Upper bound (exclusive). Defaults to 100.',
        minimum: 1,
      },
    },
  },
};

// Create a handler for the tool
async function randomIntHandler({ max = 100 }: { max?: number }) {
  return Math.floor(Math.random() * max).toString();
}

// Register the tool with AITuberOnAirCore
const aituber = new AITuberOnAirCore({
  // ... other options ...
  tools: [{ definition: randomIntTool, handler: randomIntHandler }],
});

// Set up event listeners for tool use
aituber.on(AITuberOnAirCoreEvent.TOOL_USE, (toolBlock) => 
  console.log(`Tool use -> ${toolBlock.name}`, toolBlock.input));

aituber.on(AITuberOnAirCoreEvent.TOOL_RESULT, (resultBlock) => 
  console.log(`Tool result ->`, resultBlock.content));

Tool Iteration Control

You can limit the number of tool call iterations using the maxHops option:

const aituber = new AITuberOnAirCore({
  // ... other options ...
  chatOptions: {
    systemPrompt: 'Your system prompt',
    // ... other chat options ...
    maxHops: 10,  // Maximum number of tool call iterations (default: 6)
  },
  tools: [/* your tools */],
});

Function Calling Differences

AITuber OnAir Core supports major AI providers including OpenAI, Claude, Gemini, and Z.ai. Each provider has a different implementation of function calling (tool invocation). These differences are abstracted by AITuber OnAir Core, allowing developers to use a unified interface, but understanding the background is important.

Note: This explanation covers the API versions as of May 2025. APIs are frequently updated, so please refer to the official documentation for the latest information.

OpenAI Function Calling Implementation

OpenAI's function calling has the following characteristics:

  • Tool Definition Format: Uses an array of functions (deprecated) or tools (recommended from 2023-12-01) based on JSON Schema
  • Response Format: Returns a response object containing a tool_calls array when using tools
  • Tool Result Submission: Tool results are sent as messages with role: 'tool'
  • Multiple Tool Support: Can call multiple tools simultaneously (Parallel function calling)
// OpenAI tool definition example (minimal form)
const tools = [
  {
    type: "function", 
    function: {
      name: "randomInt",
      description: "Return a random integer from 0 to (max - 1)",
      parameters: {
        type: "object",
        properties: {
          max: {
            type: "integer",
            description: "Upper bound (exclusive). Defaults to 100."
          }
        },
        required: [] // Explicitly specifying even when empty improves schema validity
      }
    }
  }
];

// OpenAI tool call response example
{
  role: "assistant",
  content: null,
  tool_calls: [
    {
      id: "call_abc123",
      type: "function",
      function: {
        name: "randomInt",
        arguments: "{\"max\":10}" // Note that this is returned as a stringified JSON
      }
    }
  ]
}

// Multiple tool calls example (Parallel function calling)
{
  role: "assistant",
  content: null,
  tool_calls: [
    {
      id: "call_abc123",
      type: "function",
      function: {
        name: "randomInt",
        arguments: "{\"max\":10}"
      }
    },
    {
      id: "call_def456",
      type: "function",
      function: {
        name: "getCurrentTime",
        arguments: "{\"timezone\":\"JST\"}"
      }
    }
  ]
}

// OpenAI tool result submission example
{
  role: "tool",
  tool_call_id: "call_abc123",
  content: "7"
}

When handling OpenAI's function calling, AITuber OnAir Core converts tool definitions to OpenAI's format and processes tool calls and results. The transformToolToFunction method in the class performs this conversion.

Claude's Tool Calling Implementation

Claude's tool calling has the following characteristics:

  • Tool Definition Format: Specifies name, description, and input_schema for each tool in the tools array
  • Response Format: Returned as a special block with type: 'tool_use' and stops with stop_reason: 'tool_use'
  • Tool Result Submission: Included in user role messages as type: 'tool_result'
  • Special Streaming Handling: Requires special logic to handle tool calls in streaming responses
// Claude tool definition example
const tools = [
  {
    name: "randomInt",
    description: "Return a random integer from 0 to (max - 1)",
    input_schema: {
      type: "object",
      properties: {
        max: {
          type: "integer",
          description: "Upper bound (exclusive). Defaults to 100."
        }
      }
    }
  }
];

// Claude tool call response example
{
  id: "msg_abc123",
  model: "claude-haiku-4-5-20251001",
  role: "assistant",
  content: [
    { type: "text", text: "I'll generate a random number for you." },
    { 
      type: "tool_use", 
      id: "tu_abc123",
      name: "randomInt",
      input: { max: 10 }
    }
  ],
  stop_reason: "tool_use"
}

// Example with only tool use, no text content
{
  id: "msg_xyz789",
  model: "claude-haiku-4-5-20251001",
  role: "assistant",
  content: [
    { 
      type: "tool_use", 
      id: "tu_xyz789",
      name: "randomInt",
      input: { max: 100 }
    }
  ],
  stop_reason: "tool_use"
}

// Claude tool result submission example
{
  role: "user",
  content: [
    {
      type: "tool_result",
      tool_use_id: "tu_abc123",
      content: "7"
    }
  ]
}

When handling Claude's tool calls, AITuber OnAir Core processes Claude's unique format and abstracts the complex processing, especially during streaming responses. Special handling is included in the runToolLoop method.

Gemini's Tool Calling Implementation

Gemini's tool calling has the following characteristics:

  • Tool Definition Format: Describes definitions in functionDeclarations within the tools array
  • Response Format: Returned as content objects containing functionCall parts
  • Tool Result Submission: Sent as functionResponse objects included in content parts
  • Compositional Calling: Supports Compositional Function Calling
// Gemini tool definition example
const tools = [
  {
    functionDeclarations: [
      {
        name: "randomInt",
        description: "Return a random integer from 0 to (max - 1)",
        parameters: {
          type: "object",
          properties: {
            max: {
              type: "integer",
              description: "Upper bound (exclusive). Defaults to 100."
            }
          }
        }
      }
    ]
  }
];

// Gemini tool call response example (note the deep structure)
{
  candidates: [
    {
      content: {
        parts: [
          {
            functionCall: {
              name: "randomInt",
              args: {
                max: 10
              }
            }
          }
        ]
      }
    }
  ]
}

// Compositional function calling example
{
  candidates: [
    {
      content: {
        parts: [
          {
            functionCall: {
              name: "randomInt",
              args: {
                max: 10
              }
            }
          },
          {
            functionCall: {
              name: "formatResult",
              args: {
                prefix: "Random number:",
                value: "<function_response:randomInt>"
              }
            }
          }
        ]
      }
    }
  ]
}

// Gemini tool result submission example
// Include functionResponse directly in content parts (SDK automatically sets the role)
{
  parts: [
    {
      functionResponse: {
        name: "randomInt",
        response: {
          value: "7"
        }
      }
    }
  ]
}

// When directly calling REST API, you might include role like this
{
  role: "function",
  parts: [
    {
      functionResponse: {
        name: "randomInt",
        response: {
          value: "7"
        }
      }
    }
  ]
}

When handling Gemini's tool calls, AITuber OnAir Core processes Gemini's complex response structure and tool result format. Special logic is needed to convert tool responses to the appropriate JSON format.

Streaming Implementation Differences

Each provider also has differences in how tool calls are processed during streaming responses:

  1. OpenAI:

    • During streaming, delta updates are sent as delta.tool_calls
    • Requires accumulation to reconstruct complete tool call data
  2. Claude:

    • SSE streaming uses special event types content_block_delta and content_block_stop
    • Sends stop_reason: "tool_use" when a tool call is completed
    • Requires a special parser to detect tool calls
  3. Gemini:

    • During streaming, functionCall may be split across chunks
    • Requires buffering to reconstruct complete JSON structures

AITuber OnAir Core abstracts these streaming processing differences, allowing you to process tool calls and results with the same interface regardless of which provider you use.

Key Differences and Abstraction Between Providers

AITuber OnAir Core abstracts the differences between these three providers and provides a unified interface:

  1. Input Format Differences:

    • Each provider uses its own tool definition format
    • AITuber OnAir Core performs appropriate conversions internally and provides a common ToolDefinition interface
  2. Response Processing Differences:

    • OpenAI uses tool_calls objects
    • Claude uses tool_use blocks
    • Gemini uses functionCall objects
    • AITuber OnAir Core processes each format and converts to unified TOOL_USE events
  3. Tool Result Submission Format Differences:

    • Each provider accepts tool results in different formats
    • AITuber OnAir Core converts and sends in the appropriate format
  4. Streaming Processing Differences:

    • Claude in particular requires special handling for tool calls during streaming
    • AITuber OnAir Core abstracts this and provides a consistent streaming experience across all providers
  5. Tool Call Iteration:

    • The runToolLoop method is implemented according to each provider's characteristics, providing consistent tool iteration

Through these abstractions, developers can use tool functionality through AITuber OnAir Core's unified interface without worrying about the details of provider implementations. Even when switching providers, there's no need to change tool definition and processing code.

Using MCP

AITuber OnAir Core allows you to integrate MCP using tool calls.

Here's an example of integration.
The following is a simple sample that integrates an MCP that returns a random number.

// mcpClient.ts
import { Client as MCPClient } from "@modelcontextprotocol/sdk/client/index.js";
import { StreamableHTTPClientTransport } from "@modelcontextprotocol/sdk/client/streamableHttp.js";
  
let clientPromise: Promise<MCPClient> | null = null;
  
async function getMcpClient(): Promise<MCPClient> {
  if (clientPromise) return clientPromise;

  const client = new MCPClient({
    name: "random-int-server",
    version: "0.0.1",
  });
  const endpoint = import.meta.env.VITE_MCP_ENDPOINT as string;
  if (!endpoint) throw new Error("VITE_MCP_ENDPOINT is not defined");

  const transport = new StreamableHTTPClientTransport(new URL(endpoint));
  clientPromise = client.connect(transport).then(() => client);
  return clientPromise;
}

export function createMcpToolHandler<T extends { [key: string]: unknown } = any>(toolName: string) {
    return async (args: T): Promise<string> => {
      const client = await getMcpClient();
      const out = await client.callTool({ name: toolName, arguments: args });
      return (out.content as { text: string }[] | undefined)?.[0]?.text ?? "";
    };
  }
import { createMcpToolHandler } from './mcpClient';

// tool definition
const randomIntTool: ToolDefinition<{ max: number }> = {
  name: 'randomInt',
  description:
    "Return a random integer from 0 (inclusive) up to, but not including, `max`. If `max` is omitted the default upper‑bound is 100.",
  parameters: {
    type: 'object',
    properties: {
      max: { type: 'integer', description: 'Exclusive upper bound for the random integer', minimum: 1 },
    },
    required: ['max'],
  },
};

// mcp tool handler
const randomIntHandler = createMcpToolHandler<{ max: number }>('randomInt');

// create options
const aituberOptions: AITuberOnAirCoreOptions = {
  chatProvider,
  apiKey: apiKey.trim(),
  model,
  chatOptions: {
    systemPrompt: systemPrompt.trim() || DEFAULT_SYSTEM_PROMPT,
    visionPrompt: visionPrompt.trim() || DEFAULT_VISION_PROMPT,
  },
  tools: [{ definition: randomIntTool, handler: randomIntHandler }],
  debug: true,
};

// create new instance
const newAITuber = new AITuberOnAirCore(aituberOptions);

Using OpenAI Remote MCP

OpenAI's Responses API allows connecting to remote MCP servers. When you specify MCP server configurations via the mcpServers option, AITuberOnAirCore automatically switches to the Responses API endpoint for OpenAI.

import {
  AITuberOnAirCore,
  AITuberOnAirCoreOptions,
  MCPServerConfig,
} from '@aituber-onair/core';

const mcpServers: MCPServerConfig[] = [
  {
    type: 'url',
    url: 'https://mcp-server.example.com/',
    name: 'example-mcp',
    require_approval: 'never', // Optional: 'always' | 'never'
    tool_configuration: { allowed_tools: ['example_tool'] },
    authorization_token: 'YOUR_TOKEN',
  },
];

const options: AITuberOnAirCoreOptions = {
  chatProvider: 'openai',
  apiKey: 'your-openai-api-key',
  model: 'gpt-4.1',
  mcpServers, // Automatically switches to Responses API when MCP servers are configured
};

const aituber = new AITuberOnAirCore(options);

Note: The endpoint configuration is OpenAI-specific and is automatically managed based on MCP server configuration. Other providers (Claude, Gemini) use their own fixed endpoints.

Using Claude MCP Connector

AITuber OnAir Core supports Claude's Model Context Protocol (MCP) connector feature, allowing you to connect to remote MCP servers directly from the Messages API without a separate MCP client.

Basic Usage

When using the Claude provider, you can specify MCP servers in the mcpServers option:

import { AITuberOnAirCore, AITuberOnAirCoreOptions } from '@aituber-onair/core';
import { MCPServerConfig } from '@aituber-onair/core';

// Define MCP server configuration
const mcpServers: MCPServerConfig[] = [
  {
    type: 'url',
    url: 'https://mcp-server.example.com/sse',
    name: 'example-mcp',
    tool_configuration: {
      enabled: true,
      allowed_tools: ['example_tool_1', 'example_tool_2']
    },
    authorization_token: 'YOUR_TOKEN' // Optional, for OAuth-enabled servers
  }
];

// Create AITuberOnAirCore instance with MCP servers
const options: AITuberOnAirCoreOptions = {
  chatProvider: 'claude', // MCP is only supported with Claude
  apiKey: 'your-claude-api-key',
  model: 'claude-haiku-4-5-20251001',
  chatOptions: {
    systemPrompt: 'You are an AI streamer with access to remote tools via MCP.',
  },
  // Traditional tools (optional, can be used alongside MCP)
  tools: [
    {
      definition: {
        name: 'local_tool',
        description: 'A local tool',
        parameters: {
          type: 'object',
          properties: {
            input: { type: 'string', description: 'Input text' }
          }
        }
      },
      handler: async (input) => {
        return `Local result: ${input.input}`;
      }
    }
  ],
  // MCP servers configuration
  mcpServers: mcpServers,
  debug: true,
};

const aituber = new AITuberOnAirCore(options);

Multiple MCP Servers

You can connect to multiple MCP servers by including multiple configurations:

const mcpServers: MCPServerConfig[] = [
  {
    type: 'url',
    url: 'https://mcp-server-1.example.com/sse',
    name: 'server-1',
    authorization_token: 'TOKEN_1'
  },
  {
    type: 'url',
    url: 'https://mcp-server-2.example.com/sse',
    name: 'server-2',
    tool_configuration: {
      enabled: true,
      allowed_tools: ['specific_tool_1', 'specific_tool_2']
    }
  }
];

OAuth Authentication

For MCP servers that require OAuth authentication, you can obtain an access token using the MCP inspector:

npx @modelcontextprotocol/inspector

Follow the OAuth flow in the inspector and copy the access_token value to use as the authorization_token in your configuration.

Event Handling

MCP tool usage is handled through the same event system as traditional tools:

// Listen for tool usage (includes both traditional tools and MCP tools)
aituber.on(AITuberOnAirCoreEvent.TOOL_USE, (toolBlocks) => {
  console.log('Tools used:', toolBlocks);
});

aituber.on(AITuberOnAirCoreEvent.TOOL_RESULT, (resultBlocks) => {
  console.log('Tool results:', resultBlocks);
});

Limitations

  • MCP connector is only available with the Claude provider
  • Only HTTP-based MCP servers are supported (STDIO servers are not supported)
  • Currently only tool calls are supported from the MCP specification
  • Not available on Amazon Bedrock and Google Vertex

Coexistence with Traditional Tools

MCP servers and traditional tool definitions can be used simultaneously. The AI can access both local tools and remote MCP tools seamlessly.

Response Length Control

AITuber OnAir Core provides comprehensive response length control functionality, allowing you to fine-tune AI response lengths for both text chat and vision processing.

Overview

Response length control helps you:

  • Optimize costs by limiting token usage
  • Control response verbosity for different scenarios
  • Maintain consistent response patterns across your application
  • Separate control for text chat and vision processing

Configuration Options

You can control response length using two approaches:

1. Direct Token Specification

Specify exact token limits directly:

const options: AITuberOnAirCoreOptions = {
  chatOptions: {
    maxTokens: 150,         // Direct token limit for text chat
    visionMaxTokens: 200,   // Direct token limit for vision processing
  },
  // ... other options
};

2. Preset Response Lengths

Use predefined presets for convenience:

const options: AITuberOnAirCoreOptions = {
  chatOptions: {
    responseLength: 'medium',        // Preset for text chat
    visionResponseLength: 'long',    // Preset for vision processing
  },
  // ... other options
};

Available presets:

  • 'veryShort': 40 tokens - Brief, essential responses only
  • 'short': 100 tokens - Concise but complete responses
  • 'medium': 200 tokens - Balanced length for most scenarios
  • 'long': 300 tokens - Detailed responses with context

Priority System

When multiple length controls are specified, the following priority order applies:

  1. Direct values (maxTokens, visionMaxTokens) - Highest priority
  2. Preset values (responseLength, visionResponseLength) - Medium priority
  3. Default values (1000 tokens) - Fallback when nothing is specified

Vision-Specific Settings

Vision processing often requires different response lengths than text chat. You can configure them separately:

const options: AITuberOnAirCoreOptions = {
  chatOptions: {
    // Text chat settings
    responseLength: 'short',      // Concise text responses
    
    // Vision processing settings
    visionResponseLength: 'long', // Detailed image descriptions
  },
};

If vision-specific settings are not provided, they will fall back to the regular chat settings.

Dynamic Updates

Response length settings can be updated at runtime:

// Update chat processor options
aituber.updateChatProcessorOptions({
  maxTokens: 100,
  visionMaxTokens: 250,
});

Usage Examples

Different Lengths for Different Scenarios

// Short responses for quick interactions
const quickChat = new AITuberOnAirCore({
  chatOptions: {
    responseLength: 'veryShort',
  },
});

// Detailed responses for educational content
const educationalChat = new AITuberOnAirCore({
  chatOptions: {
    responseLength: 'long',
    visionResponseLength: 'long',
  },
});

Mixing Direct Values and Presets

const options: AITuberOnAirCoreOptions = {
  chatOptions: {
    maxTokens: 120,                  // Direct value for text chat
    visionResponseLength: 'medium',  // Preset for vision
  },
};

Provider Compatibility

Response length control is supported across all AI providers:

  • OpenAI: Uses max_tokens parameter
  • Claude: Uses max_tokens parameter
  • Gemini: Uses maxOutputTokens parameter

The implementation handles provider-specific differences automatically.

Architecture

AITuberOnAirCore is designed with the following layered structure:

AITuberOnAirCore (Integration Layer)
    ├── ChatProcessor (Conversation handling)
    │     └── ChatService (AI Chat)
    ├── MemoryManager (Memory handling)
    │     └── Summarizer (Summarization)
    └── VoiceService (Speech processing)
          └── VoiceEngineAdapter (Speech Engine Interface)
                └── Various Speech Engines (VOICEVOX, OpenAI, etc.)

Directory Structure

The source code is organized around the following directory structure:

src/
  ├── constants/             # Constants and configuration
  │     ├── index.ts         # Exported constants
  │     └── prompts.ts       # Default prompts and templates
  ├── core/                  # Core components
  │     ├── AITuberOnAirCore.ts
  │     ├── ChatProcessor.ts
  │     └── MemoryManager.ts
  ├── services/              # Service implementations
  │     ├── chat/            # Chat services
  │     │    ├── ChatService.ts            # Base interface
  │     │    ├── ChatServiceFactory.ts     # Factory for providers
  │     │    └── providers/                # AI provider implementations
  │     │         ├── ChatServiceProvider.ts  # Provider interface
  │     │         ├── claude/              # Claude-specific
  │     │         │    ├── ClaudeChatService.ts
  │     │         │    ├── ClaudeChatServiceProvider.ts
  │     │         │    └── ClaudeSummarizer.ts
  │     │         ├── gemini/              # Gemini-specific
  │     │         │    ├── GeminiChatService.ts
  │     │         │    ├── GeminiChatServiceProvider.ts
  │     │         │    └── GeminiSummarizer.ts
  │     │         └── openai/              # OpenAI-specific
  │     │              ├── OpenAIChatService.ts
  │     │              ├── OpenAIChatServiceProvider.ts
  │     │              └── OpenAISummarizer.ts
  │     ├── voice/           # Voice services
  │     │    ├── VoiceService.ts
  │     │    ├── VoiceEngineAdapter.ts
  │     │    └── engines/    # Voice engine implementations
  │     └── youtube/         # YouTube API integration
  │          └── YouTubeDataApiService.ts  # YouTube Data API client
  ├── types/                 # TypeScript type definitions
  └── utils/                 # Utilities and helpers
       ├── screenplay.ts     # Text and emotion processing
       └── storage.ts        # Storage utilities

Main Components

AITuberOnAirCore

This is the overall integration class, responsible for initializing and coordinating other components. It extends EventEmitter and emits events at various processing stages. In most cases, you will interact primarily with this class to use its features.

Main methods include:

  • processChat(text) – Process text input
  • processVisionChat(imageDataUrl, visionPrompt?) – Process image input (optionally pass a custom prompt)
  • stopSpeech() – Stop speech playback
  • getChatHistory() – Retrieve chat history
  • setChatHistory(messages) – Set chat history from external source (e.g., for replay or migration)
  • clearChatHistory() – Clear chat history
  • updateVoiceService(options) – Update speech settings
  • isMemoryEnabled() – Check if memory functionality is enabled
  • generateOneShotContentFromHistory(prompt, messageHistory) – Generate new content from a system prompt and provided message history (one-shot, no impact on internal chat history)
  • offAll() – Remove all event listeners

ChatProcessor

The component that sends text input to an AI model (e.g., OpenAI GPT) and receives responses. It manages the conversation flow and supports streaming responses. It also handles emotion extraction from responses.

  • updateOptions(newOptions) – Allows you to update settings at runtime

MemoryManager

MemoryManager is designed to prevent issues such as API token limits, increased costs, and slow responses that can occur when the chat log grows too large. When a certain time or message threshold is exceeded, older chat history is summarized and stored as short-, mid-, and long-term memory. This allows recent conversation to be sent as-is, while past context is provided as a summary, maintaining context for the AI while keeping API requests efficient.

Handles conversational context. In long conversations, older messages are summarized and maintained as short-term (1 min), mid-term (4 min), and long-term (9 min) memory. This helps maintain consistency in AI responses.

  • Custom Settings:
    • summaryPromptTemplate can be customized for summarization (it uses a {maxLength} placeholder).

VoiceService

Converts text to speech. It integrates with multiple external speech synthesis engines through the VoiceEngineAdapter.

speakTextWithOptions Method

The AITuberOnAirCore class provides a flexible speakTextWithOptions method for speech playback:

// Example of speaking text with temporary settings
await aituberOnairCore.speakTextWithOptions('[happy] Hello, everyone watching!', {
  // Enable or disable avatar animation
  enableAnimation: true,
  
  // Temporarily override current speech settings
  temporaryVoiceOptions: {
    engineType: 'voicevox',
    speaker: '8',
    apiKey: 'YOUR_API_KEY'  // If required
  },
  
  // Specify the ID of the HTML audio element for playback
  audioElementId: 'custom-audio-player'
});

Key Features:

  1. Temporary Voice Settings: Override current speech settings without permanently changing them.
  2. Animation Control: Control avatar animation with the enableAnimation option.
  3. Flexible Audio Playback: Play audio in a specified HTML audio element.
  4. Automatic Emotion Extraction: Extract emotion tags (e.g., [happy]) from text and provide them in the SPEECH_START event.

Event System

AITuberOnAirCore emits the following events:

  • PROCESSING_START: When processing begins
  • PROCESSING_END: When processing finishes
  • ASSISTANT_PARTIAL: Upon receiving partial responses from the assistant (streaming)
  • ASSISTANT_RESPONSE: Upon receiving a complete response (includes a screenplay object and rawText with emotion tags)
  • SPEECH_START: When speech playback starts (includes a screenplay object with emotion and rawText with emotion tags)
  • SPEECH_END: When speech playback ends
  • TOOL_USE: When the AI calls a tool (includes the name of the tool and its input parameters)
  • TOOL_RESULT: When a tool execution completes and returns a result
  • ERROR: When an error occurs
  • CHAT_HISTORY_SET: When chat history is set
  • CHAT_HISTORY_CLEARED: When chat history is cleared
  • MEMORY_CREATED: When a new memory is created
  • MEMORY_REMOVED: When memory is removed
  • MEMORY_LOADED: When memory is loaded from storage
  • MEMORY_SAVED: When memory is saved to storage
  • STORAGE_CLEARED: When storage is cleared

Safely Handling Event Data

In particular, when implementing a listener for the SPEECH_START event, it is recommended to check if data is present:

// Safe handling of SPEECH events
aituber.on(AITuberOnAirCoreEvent.SPEECH_START, (data) => {
  if (!data) {
    console.log('No data available');
    return;
  }
  
  const screenplay = data.screenplay;
  if (!screenplay) {
    console.log('No screenplay object');
    return;
  }
  
  const emotion = screenplay.emotion || 'neutral';
  console.log(`Speech started: Emotion = ${emotion}`);
  
  // Get original text with emotion tags
  console.log(`Original text: ${data.rawText}`);
  
  // Update UI or avatar animation
  updateUIWithEmotion(emotion);
});

Emotion Handling

In a React application, you might use useRef to store the latest emotion data for immediate access:

// Example in a React component
const [currentEmotion, setCurrentEmotion] = useState('neutral');
const emotionRef = useRef({ emotion: 'neutral', text: '' });

useEffect(() => {
  if (aituberOnairCore) {
    aituberOnairCore.on(AITuberOnAirCoreEvent.SPEECH_START, (data) => {
      if (data?.screenplay?.emotion) {
        setCurrentEmotion(data.screenplay.emotion);
        emotionRef.current = data.screenplay;
      }
    });
  }
}, [aituberOnairCore]);

// Use the ref for animation callbacks
const handleAnimation = () => {
  const emotion = emotionRef.current.emotion || 'neutral';
  // Perform animation based on emotion
};

ChatProcessor Events

The internal ChatProcessor emits additional events:

  • chatLogUpdated: Fired when the chat log is updated (e.g., when new messages are added or history is cleared).

You can access this event by referencing the ChatProcessor instance directly:

// Example: using the chatLogUpdated event in ChatProcessor
const aituber = new AITuberOnAirCore(options);
const chatProcessor = aituber['chatProcessor']; // Accessing internal component

chatProcessor.on('chatLogUpdated', (chatLog) => {
  console.log('Chat log updated:', chatLog);
  
  // Example: Update UI
  updateChatDisplay(chatLog);
  
  // Example: Sync with an external system
  syncChatToExternalSystem(chatLog);
});

Possible use cases for chatLogUpdated include:

  1. Real-Time Chat UI Updates
    Reflect new messages or cleared logs in the UI immediately.
  2. External System Integration
    Save chat logs to a database or send them to an analytics service.
  3. Debugging & Monitoring
    Monitor changes in the chat log during development.

Supported Speech Engines

AITuberOnAirCore supports the following speech engines:

  • VOICEVOX: High-quality Japanese speech synthesis engine.
  • VoicePeak: Speech synthesis engine with rich emotional expression, supporting single-tag or weighted emotion overrides.
  • AivisSpeech: Speech synthesis using AI technology.
  • Aivis Cloud: High-quality Japanese text-to-speech service with SSML support, emotional intensity control, and multiple output formats (WAV, FLAC, MP3, AAC, Opus).
  • OpenAI TTS: Text-to-speech API from OpenAI.
  • Gemini TTS: Gemini API-based text-to-speech with selectable preview TTS models including gemini-3.1-flash-tts-preview, plus style/audio-tag prompt support.
  • xAI TTS: xAI text-to-speech with selectable codec, sample rate, and bit rate options.
  • Unreal Speech: Unreal Speech v8 /stream endpoint with bitrate, speed, pitch, codec, and temperature options.
  • ElevenLabs: ElevenLabs Text to Speech API with model, output format, language code, voice settings, and text normalization options.
  • Fish Audio: Fish Audio one-shot TTS with S2 Pro by default, configurable output/latency options, and reference-voice list discovery.
  • Cartesia: Cartesia synchronous TTS with Sonic 3.5 / 3.6, language/output controls, and voice-list discovery.
  • Deepgram Flux: English-only, one-shot MP3 speech with a public voice catalog and optional speed control.
  • Inworld: Inworld TTS REST API with selectable model, audio encoding, sample rate, bitrate, language, delivery mode, and temperature options.
  • Gradium: Gradium REST TTS API with selectable preset voices, output format, temperature, similarity, padding, and rewrite-rule options.
  • OpenAI-Compatible TTS: Self-hosted or third-party /v1/audio/speech compatible endpoints.
  • MiniMax: Multi-language TTS using the current T2A v2 models and documented system voice presets. An API key is required; GroupId is a legacy optional query parameter.
  • Piper Plus: Browser WASM TTS using ONNX Runtime Web and OpenJTalk assets for on-device synthesis.
  • Web Speech API: Browser-native speech synthesis with runtime voice-list discovery and configurable rate, pitch, volume, and language. It plays audio directly through the browser and does not expose audio bytes, so the Core React examples do not support audio-buffer-based lip sync with this engine.
  • None: No voice mode (no audio output).

You can dynamically switch the speech engine via updateVoiceService:

// Example of switching speech engines
aituber.updateVoiceService({
  engineType: 'openai',
  speaker: 'alloy',
  apiKey: 'YOUR_OPENAI_API_KEY'
});

Speech Chunking (Optional)

Speech synthesis can optionally split assistant responses into smaller chunks so that playback starts sooner for long messages. Chunking is disabled by default to preserve the previous behaviour (one TTS request per response).

You can enable and configure it when creating AITuberOnAirCore:

const aituber = new AITuberOnAirCore({
  // ... existing options ...
  speechChunking: {
    enabled: true,
    // Minimum "word" count before a new chunk is started. Japanese text falls
    // back to character counts when spaces are not present.
    minWords: 40,
    // Pick a preset for punctuation detection (ja | en | ko | zh | all)
    locale: 'ja',
    // Or override with your own separator characters
    // separators: ['.', '!', '?'],
  },
});
  • When enabled, the core splits text at punctuation (。!?!? and commas) and merges adjacent segments until the minWords threshold is reached. This keeps the number of TTS requests manageable while still reducing perceived latency.
  • Setting minWords to 0 (or omitting it) keeps the raw punctuation-based segments and simply streams them in order.
  • You can toggle the behaviour at runtime, for example when the viewer switches between local and cloud TTS engines:
aituber.updateSpeechChunking({ enabled: false });
aituber.updateSpeechChunking({
  enabled: true,
  minWords: 25,
  locale: 'en',
  separators: ['.', '!', '?'],
});

Custom API Endpoints

For locally hosted or overridable voice engines (VOICEVOX, VoicePeak, AivisSpeech, OpenAI-Compatible TTS, Unreal Speech, ElevenLabs, Inworld, Gradium), you can specify custom API endpoint URLs:

// Example of setting custom API endpoints
aituber.updateVoiceService({
  engineType: 'voicevox',
  speaker: '1',
  // Custom endpoint for a self-hosted or alternative VOICEVOX server
  voicevoxApiUrl: 'http://custom-voicevox-server:50021'
});

// Example for VoicePeak
aituber.updateVoiceService({
  engineType: 'voicepeak',
  speaker: '2',
  voicepeakApiUrl: 'http://custom-voicepeak-server:20202',
  voicepeakEmotion: { happy: 40, fun: 60 },
  voicepeakSpeed: 140,
  voicepeakPitch: 20,
});

// Example for AivisSpeech
aituber.updateVoiceService({
  engineType: 'aivisSpeech',
  speaker: '3',
  aivisSpeechApiUrl: 'http://custom-aivis-server:10101'
});

// Example for OpenAI-compatible TTS
aituber.updateVoiceService({
  engineType: 'openaiCompatible',
  openAiCompatibleApiUrl: 'http://localhost:8880/v1/audio/speech',
  openAiCompatibleModel: 'your-model-id',
  openAiCompatibleSpeed: 1.1
});

// Example for Aivis Cloud (high-quality Japanese TTS with SSML support)
aituber.updateVoiceService({
  engineType: 'aivisCloud',
  speaker: 'YOUR_SPEAKER_UUID', // Speaker UUID from Aivis Cloud
  apiKey: 'YOUR_AIVIS_CLOUD_API_KEY',
  // Optional parameters for advanced control
  emotionalIntensity: 1.0,     // 0.0-2.0 range for emotional expression
  speakingRate: 1.0,           // 0.5-2.0 range for speaking speed
  outputFormat: 'wav'          // wav, flac, mp3, aac, opus
});

// Example for MiniMax (simplified configuration)
aituber.updateVoiceService({
  engineType: 'minimax',
  speaker: 'Japanese_IntellectualSenior', // or another system voice preset
  apiKey: 'YOUR_MINIMAX_API_KEY',
  groupId: 'YOUR_GROUP_ID',  // Legacy optional query parameter
  endpoint: 'global'         // Optional: 'global' (default) or 'china'
});

// Example for Unreal Speech
aituber.updateVoiceService({
  engineType: 'unrealSpeech',
  speaker: 'af_bella',
  apiKey: 'YOUR_UNREAL_SPEECH_API_KEY',
  unrealSpeechBitrate: '192k',
  unrealSpeechCodec: 'libmp3lame',
  unrealSpeechSpeed: 0.3,
});

// Example for ElevenLabs
aituber.updateVoiceService({
  engineType: 'elevenLabs',
  speaker: 'YOUR_ELEVENLABS_VOICE_ID',
  apiKey: 'YOUR_ELEVENLABS_API_KEY',
  elevenLabsModel: 'eleven_flash_v2_5',
  elevenLabsOutputFormat: 'mp3_44100_128',
  elevenLabsLanguageCode: 'ja',
});

// MiniMax GroupId is a legacy optional query parameter.
// MiniMax also supports region-specific endpoints:
// - 'global': For international users (default)
// - 'china': For users in mainland China

This is useful when running voice engines on different ports or remote servers.

AI Provider System

AITuber OnAir Core adopts an extensible provider system, enabling integration with various AI APIs. Currently, OpenAI API, OpenAI-compatible APIs, Gemini API, Gemini Nano (Chrome Built-in AI), Claude API, xAI API, Z.ai API, Kimi API, and OpenRouter API, DeepSeek API, Mistral API, Sakana AI, and PLaMo are available. If you would like to use any other API, please submit a PR or send us a message.

Available Providers

Currently, the following AI provider is built-in:

  • OpenAI: Supports models like GPT-5 family (Nano/Mini/Standard/5.1/5.4/5.5/5.6 Sol/Terra/Luna/5.4 Mini/5.4 Nano/5.4 Pro), GPT-4.1 (including Mini/Nano), GPT-4o, GPT-4o-mini, O3-mini, o1, o1-mini. GPT-5.6 models also support max reasoning effort.
  • Gemini: Supports models like Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash / Flash-Lite, Gemini 3.1 Flash-Lite, Gemini 3.1 Pro Preview, Gemini 3 Flash Preview, Gemini 2.5 Pro, Gemini 2.5 Flash, Gemini 2.5 Flash Lite, Gemma 4 31B IT, and Gemma 4 26B A4B IT. Gemini 3 Flash-family models support configurable reasoning_effort; Gemini 3.8 Flash and Gemini 3.7 Flash start at low, while earlier Flash models start at minimal and Pro models at low. Gemini 3.8 Flash also supports Vision and tool calling through the native Gemini provider. Gemini 2.5 continues to use thinkingBudget instead.
  • Gemini Nano: Supports the built-in Chrome gemini-nano model without an API key (Chrome 138+ with Prompt API flags enabled)
  • Claude: Supports current Claude API model IDs including Claude Fable 5, Claude Opus 5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Opus 4.5, Claude Sonnet 4.6, Claude Sonnet 4.5, and Claude Haiku 4.5. Retired IDs remain exported for source compatibility but are not advertised in selectors. Supported models accept reasoning_effort, which maps to Anthropic's output_config.effort; the API default is high.
  • xAI: Supports Grok 4.6, Grok 4.5, Grok 4.3, and Grok 4.20 model families. Grok 4.6 supports low, medium, high, and xhigh reasoning effort and defaults to low; Grok 4.5 also defaults to low, while Grok 4.3 defaults to none for lower latency. Retired Grok 4.1 Fast IDs remain compatibility exports only.
  • DeepSeek: Supports DeepSeek V4 Flash, DeepSeek V4 Pro, and the vision-capable DeepSeek V4 Flash Vision Exp through the first-class deepseek provider. These models expose model-aware reasoning_effort; Core keeps Chat's low-latency none default, while higher supported efforts remain selectable. Thinking and tool calling cannot currently be combined in one request.
  • Mistral: Supports the vision-capable Ministral 3 family (ministral-3b-2512, ministral-8b-2512, ministral-14b-2512) and current Mistral generalist models such as mistral-small-latest, mistral-medium-3-5, and mistral-large-latest, including adjustable reasoning for supported models.
  • Sakana AI: Supports Fugu (fugu), the exact Fugu Ultra v1.1 ID (fugu-ultra-v1.1), and the Japanese-focused, vision-capable Sakana Namazu (sakana-namazu) through the first-class sakana provider. Namazu disables optional thinking by default for responsive chat and can enable it explicitly through chat template options. Older aliases remain compatibility exports. Browser examples show this provider as disabled because direct browser requests can fail with CORS; use Node.js or a backend proxy.
  • PLaMo: Supports PLaMo 3.0 Prime (plamo-3.0-prime) through the first-class plamo provider. PLaMo 2.2 Prime remains a deprecated compatibility export ahead of its scheduled retirement.
  • Z.ai: Supports GLM-5.3, the vision-capable GLM-5.3 Flash, GLM-5.2, GLM-5.1, GLM-5/GLM-5-Turbo (text-only), GLM-5V-Turbo (vision), GLM-4.7, GLM-4.7 Flash/FlashX, GLM-4.6, GLM-4.6V, and GLM-4.6V Flash/FlashX. GLM-5.3 models always use thinking, expose only low, high, and max, and default to the minimum low level for responsive chat. GLM-5.2 continues to default to none.
  • Kimi: Supports Kimi K3 (kimi-k3), Kimi K2.7 Code (kimi-k2.7-code), Kimi K2.7 Code HighSpeed (kimi-k2.7-code-highspeed), Kimi K2.6 (kimi-k2.6), and Kimi K2.5 (kimi-k2.5) with vision support. Models that expose reasoning_effort use their documented supported values and defaults: Kimi K3 accepts low, high, and max, defaults to max, and cannot disable reasoning. Kimi K2.7 Code models require thinking mode.
  • OpenRouter: Supports a curated OpenRouter model list, including Auto Router and Auto Router Beta, Fusion, latest-family aliases, OpenAI GPT-5.6, Claude Fable 5 / Sonnet 5 / Opus 5 / Opus 4.8, Gemini 3.7/3.6/3.5, Z.ai GLM-5.3 / GLM-5.3 Flash / GLM-5.2, xAI Grok 4.6/4.5, Kimi K3 / Kimi K2.6, Qwen3.8 Flash, KAT-Coder V2.5, and DeepSeek V4 Flash / V4 Flash Vision Exp / V4 Pro 0813. Model-aware vision and reasoning controls use Chat's current per-model values and low-latency defaults; catalog-absent IDs remain compatibility exports but are not advertised.
  • OpenAI-Compatible: Supports arbitrary OpenAI-compatible Chat Completions endpoints; vision capability is treated as unknown until the target endpoint/model responds

For OpenRouter free-tier discovery, you can also use refreshOpenRouterFreeModels via @aituber-onair/core (re-exported from @aituber-onair/chat).

Specifying a Provider

You can specify the provider when instantiating AITuberOnAirCore:

const aituberCore = new AITuberOnAirCore({
  chatProvider: 'openai',  // Provider name
  apiKey: 'your-api-key',
  model: 'gpt-4o-mini',    // Optional (if omitted, the default model 'gpt-4o-mini' will be used)
  // Other options...
});

Model-Specific Feature Limitations

Different AI models support different features. For example:

  • GPT-4o, GPT-4o-mini: Support both text chat and image processing (Vision)
  • O3-mini: Supports text chat only (does not support image processing)
  • GPT-5.4 Pro: Uses Responses API only (Chat Completions API is not available)
  • GPT-5.5 Pro: Not included in supported models because OpenAI documents it as non-streaming, while the standard chat flow expects streaming support

When selecting a model, be aware of these limitations. Attempting to use unsupported features will result in an explicit error.

For openai-compatible, vision support is treated as unknown because the library cannot pre-validate arbitrary local/self-hosted endpoints. Image requests are allowed, but unsupported endpoint/model combinations will fail at runtime with the upstream API error.

Note: If you don't specify a model, the default model used is 'gpt-4o-mini'. This model supports both text chat and image processing.

Using Different Models Together

If you want to use different models for text chat and image processing, you can use the visionModel option:

const aituberCore = new AITuberOnAirCore({
  apiKey: 'your-api-key',
  chatProvider: 'openai',
  model: 'o3-mini',       // For text chat 
  visionModel: 'gpt-4o',  // For image processing
  // Other options...
});

This allows for optimizations such as using a lightweight model for text chat and a more powerful model only when image processing is needed.

Note: When specifying a visionModel, ensure it supports vision capabilities. Built-in providers are validated during initialization. For openai-compatible, support is treated as unknown, so the request is allowed and the endpoint may reject it at runtime if vision is unsupported.

If you need to surface this in your own UI, @aituber-onair/core re-exports the chat capability helpers:

import { ChatServiceFactory, type VisionSupportLevel } from '@aituber-onair/core';

const visionSupport: VisionSupportLevel =
  ChatServiceFactory.getVisionSupportLevelForModel(
    'openai-compatible',
    'local-model',
  );
// 'supported' | 'unsupported' | 'unknown'

Retrieving Providers & Models

You can programmatically retrieve available providers and their supported models:

// Get all available providers
const providers = AITuberOnAirCore.getAvailableProviders();

// Get supported models for a specific provider
const models = AITuberOnAirCore.getSupportedModels('openai');

Creating a Custom Provider

To add a new AI provider, implement the ChatServiceProvider interface in a custom class and register it with the ChatServiceFactory:

import { ChatServiceFactory } from 'aituber-onair-core';
import { MyCustomProvider } from './MyCustomProvider';

// Register the custom provider
ChatServiceFactory.registerProvider(new MyCustomProvider());

// Use the registered provider
const aituberCore = new AITuberOnAirCore({
  chatProvider: 'myCustomProvider',
  apiKey: 'your-api-key',
  // Other options...
});

Agent SDK providers (Codex / Claude Agent SDK / Copilot)

Agent SDK providers use local subscription authentication instead of API keys. They are available from the Node.js-only @aituber-onair/core/agent entry and are kept out of the main entry so browser applications do not load or register them.

The ./agent subpath is resolved through package exports. TypeScript consumers must use moduleResolution: "node16", "nodenext", or "bundler"; legacy "node" resolution cannot resolve it. This is the same requirement as @aituber-onair/chat/agent.

Install only the SDK package that matches the provider your application uses:

npm install @aituber-onair/core @openai/codex-sdk
# or
npm install @aituber-onair/core @anthropic-ai/claude-agent-sdk
# or
npm install @aituber-onair/core @github/copilot-sdk

The current provider-to-package mappings are codex-sdk to @openai/codex-sdk, claude-agent-sdk to @anthropic-ai/claude-agent-sdk, and copilot-sdk to @github/copilot-sdk. These SDKs are loaded dynamically and are not dependencies of @aituber-onair/core.

Use createAgentChatService when you only need a standalone chat service:

import { createAgentChatService } from '@aituber-onair/core/agent';

const service = createAgentChatService('codex-sdk', {
  workingDirectory: process.cwd(),
  skipGitRepoCheck: true,
});

const result = await service.chatOnce(
  [{ role: 'user', content: 'Give me one short news headline.' }],
  false,
);

The same entry registers the providers for AITuberOnAirCore, so the full Core flow does not need an API key:

import { AITuberOnAirCore } from '@aituber-onair/core/agent';

const core = new AITuberOnAirCore({
  chatProvider: 'codex-sdk',
  chatOptions: {
    systemPrompt: 'You are a concise news presenter.',
  },
  providerOptions: {
    workingDirectory: process.cwd(),
    skipGitRepoCheck: true,
  },
  voiceOptions: {
    engineType: 'aivisSpeech',
    speaker: '888753760',
  },
});

await core.processChat('What should we cover today?');

Current limitations apply to the Agent SDK provider group as a whole: it is Node.js-only and text-only, and Core tools, vision chat, and MCP servers are not supported. Streaming capability follows the underlying chat provider metadata; currently Codex SDK returns a completed response, while Claude Agent SDK and Copilot SDK can forward partial text when their SDKs emit it. Core memory summarization is not supported with Agent SDK providers; enabling memoryOptions.enableSummarization throws a clear initialization error. Missing SDK packages or unavailable local authentication are reported at runtime with the underlying SDK error details.

Memory & Persistence

AITuberOnAirCore includes a memory feature that maintains the context of long-running conversations. The AI summarizes older messages, preserving short-, mid-, and long-term context for more coherent responses.

Memory Types

There are three types of memory:

  1. Short-Term Memory

    • Generated 1 minute after the conversation starts
    • Holds recent conversation details
  2. Mid-Term Memory

    • Generated 4 minutes after the conversation starts
    • Holds slightly broader summaries of the conversation
  3. Long-Term Memory

    • Generated 9 minutes after the conversation starts
    • Holds key themes and important information from the overall conversation

These memory records are automatically included in the AI prompts, helping the AI respond consistently over time.

Memory Persistence

AITuberOnAirCore has a pluggable design for memory persistence, so that the conversation context can be retained even if the application is restarted.

MemoryStorage Interface

Persistence is provided through the abstract MemoryStorage interface:

interface MemoryStorage {
  load(): Promise<MemoryRecord[]>;
  save(records: MemoryRecord[]): Promise<void>;