genai-lite
v0.19.0
Published
A lightweight, portable toolkit for interacting with various Generative AI APIs.
Downloads
211
Readme
genai-lite
A lightweight, portable Node.js/TypeScript library providing a unified interface for interacting with multiple Generative AI providers—both cloud-based (OpenAI, Anthropic, Google Gemini, Mistral, OpenRouter) and local (llama.cpp, stable-diffusion.cpp). Supports both LLM chat and AI image generation.
Features
🔌 Unified API - Single interface for multiple AI providers
🏠 Local & Cloud Models - Run models locally with llama.cpp or use cloud APIs
⚡ Text Streaming - Async iterable token deltas for OpenAI, Anthropic, Gemini, Mistral, OpenRouter, and llama.cpp
🖼️ Image Generation - First-class support for AI image generation (OpenAI, local diffusion)
🔐 Flexible API Key Management - Bring your own key storage solution
📦 Zero Electron Dependencies - Works in any Node.js environment
🎯 TypeScript First - Full type safety and IntelliSense support
⚡ Lightweight - Minimal dependencies, focused functionality
🛡️ Provider Normalization - Consistent responses across different AI APIs
🧠 Local Reasoning Toggle - Turn thinking on/off for detected GGUF models (Qwen 3.5-class, Gemma 4, and more) with vendor-tuned sampling defaults
🔁 Built-in Reliability - Automatic retries with backoff/
Retry-After, per-request timeouts, and abort signals — for both LLM and image generation (local diffusion additionally gets server-side cancel and is never blind-retried)🎨 Configurable Model Presets - Built-in presets with full customization options
🎭 Template Engine - Sophisticated templating with conditionals and variable substitution
📊 Configurable Logging - Debug mode, custom loggers (pino, winston), and silent mode for tests
Extensible Content Tokenizers - Exact built-ins plus append-only local tokenizer recipes and aliases
Prepared calls can inspect and budget the immutable semantic provider request, then dispatch that same representation with truthful token/termination evidence.
Installation
npm install genai-liteSet API keys as environment variables:
export OPENAI_API_KEY=sk-...
export ANTHROPIC_API_KEY=sk-ant-...
export GEMINI_API_KEY=AIza...Quick Start
Cloud Providers (OpenAI, Anthropic, Gemini, Mistral)
import { LLMService, fromEnvironment } from 'genai-lite';
const llmService = new LLMService(fromEnvironment);
const response = await llmService.sendMessage({
providerId: 'openai',
modelId: 'gpt-4.1-mini',
messages: [
{ role: 'system', content: 'You are a helpful assistant.' },
{ role: 'user', content: 'Hello, how are you?' }
]
});
if (response.object === 'chat.completion') {
console.log(response.choices[0].message.content);
}Local Models (llama.cpp)
import { LLMService } from 'genai-lite';
// Start llama.cpp server first: llama-server -m /path/to/model.gguf --port 8080
const llmService = new LLMService(async () => 'not-needed');
const response = await llmService.sendMessage({
providerId: 'llamacpp',
modelId: 'llamacpp', // Generic ID for whatever model is loaded
messages: [
{ role: 'system', content: 'You are a helpful assistant.' },
{ role: 'user', content: 'Explain quantum computing briefly.' }
]
});
if (response.object === 'chat.completion') {
console.log(response.choices[0].message.content);
}Streaming Text
import { LLMService } from 'genai-lite';
const llmService = new LLMService(async () => 'not-needed');
for await (const event of llmService.streamMessage({
providerId: 'llamacpp',
modelId: 'llamacpp',
messages: [{ role: 'user', content: 'Say hello in five words.' }],
settings: { maxTokens: 32 }
})) {
if (event.type === 'content_delta') {
process.stdout.write(event.delta);
} else if (event.type === 'reasoning_delta') {
process.stderr.write(event.delta);
} else if (event.type === 'complete') {
console.log('\nTokens:', event.response.usage?.total_tokens);
} else if (event.type === 'error') {
console.error(event.error.error.message);
}
}Streaming is implemented for all text providers: openai, anthropic, gemini, mistral, openrouter, and llamacpp. The final complete.response event contains the same normalized response shape returned by sendMessage().
Prepared Calls
const prepared = await llmService.prepareMessage(request, { mode: 'complete' });
if ('object' in prepared) throw new Error(prepared.error.message);
const inspection = await llmService.inspectPrepared(prepared);
if ('object' in inspection) throw new Error(inspection.error.message);
console.log(inspection.promptAccounting, inspection.outputTokenLimit);
const response = await llmService.sendPrepared(prepared);Preparation is credential-free and mode-bound. See Prepared Calls & Token Accounting for certified structural bounds, the single-margin capacity formula, response evidence with independent raw-content and provider-output scopes, exact active-template llama.cpp counting, and optional authoritative endpoint-generation bindings. Hosts that can guarantee revision changes for every model/build/template change may opt into preparation-state reuse; live dispatch validation still rejects stale calls.
Optional Content Tokenizers
Applications can load a pinned tokenizer recipe during startup or when a model is installed, register exact provider/model aliases, and obtain synchronous model-quality content counts:
npm install @huggingface/tokenizers@^0.1.3import { registerContentTokenProfileConfiguration } from "genai-lite";
import { loadContentTokenizerProfile } from "genai-lite/tokenizer-loader";
import {
GEMMA_4_IT_CONTENT_TOKENIZER_RECIPE,
} from "genai-lite/tokenizer-recipes";
const backend = await loadContentTokenizerProfile(
GEMMA_4_IT_CONTENT_TOKENIZER_RECIPE,
{ cacheDir: "./tokenizer-cache", allowDownload: true }
);
registerContentTokenProfileConfiguration({
backends: [backend],
aliases: [{
providerId: "llamacpp",
modelId: "gemma-4-12b-it-IQ4_XS.gguf",
profileId: backend.id,
}],
});Bundled ESM applications can statically import the optional peer and pass
tokenizersPeer: { module: tokenizersModule, packageVersion: "0.1.3" } to
loadContentTokenizerProfile(). This avoids loader-relative runtime resolution;
the asserted version is validated and recorded in runtime provenance. See the
guide below for the complete injection and bundler contract.
Registration is process-global, synchronous, transactional, and append-only. Reads do not close the registry, but existing backend IDs and aliases cannot be replaced. Re-query a previously unavailable resolution or capability after a successful addition. Registered profiles remain model evidence and cannot enter certificate APIs. See Prepared Calls & Token Accounting for cache integrity, exact aliasing, bundler configuration, and trust levels.
Image Generation
import { ImageService, fromEnvironment } from 'genai-lite';
const imageService = new ImageService(fromEnvironment);
const result = await imageService.generateImage({
providerId: 'openai-images',
modelId: 'gpt-image-1-mini',
prompt: 'A serene mountain lake at sunrise, photorealistic',
settings: {
width: 1024,
height: 1024,
quality: 'high'
}
});
if (result.object === 'image.result') {
require('fs').writeFileSync('output.png', result.data[0].data);
}Documentation
Comprehensive documentation is available in the genai-lite-docs folder.
Getting Started
- Documentation Hub - Navigation and overview
- Core Concepts - API keys, presets, settings, errors
API Reference
- LLM Service - Text generation and chat
- Prepared Calls & Token Accounting - Inspect, budget, and dispatch one canonical request
- Image Service - Image generation (cloud and local)
- llama.cpp Integration - Local LLM inference
Utilities & Advanced
- Prompting Utilities - Template engine, token counting, content parsing
- Constrained Answer Labels - Label grammars, probability evidence, and optional shared-prefix resolution
- Logging - Configure logging and debugging
- TypeScript Reference - Type definitions
Provider Reference
- Providers & Models - Supported providers and models
Examples & Help
- Example: Chat Demo - Reference implementation for chat applications
- Example: Image Demo - Reference implementation for image generation applications
- Troubleshooting - Common issues and solutions
Supported Providers
LLM Providers
- OpenAI - GPT-5 (5.2, 5.1, mini, nano), GPT-4.1, o4-mini
- Anthropic - Claude 4.5 (Opus, Sonnet, Haiku), Claude 4, Claude 3.7, Claude 3.5
- Google Gemini - Gemini 3 (Pro, Flash preview), Gemini 2.5, Gemma 3 & 4 (free)
- Mistral - Codestral, Devstral
- OpenRouter - Unified gateway to 100+ models (unknown models assumed reasoning-capable)
- llama.cpp - Run any GGUF model locally (no API keys required); streaming and local reasoning toggle for detected models
Image Providers
- OpenAI Images - gpt-image-1, dall-e-3, dall-e-2
- genai-electron - Local Stable Diffusion models
See Providers & Models for complete model listings and capabilities.
API Key Management
genai-lite uses a flexible API key provider pattern. Use the built-in fromEnvironment provider or create your own:
import { ApiKeyProvider, LLMService } from 'genai-lite';
const myKeyProvider: ApiKeyProvider = async (providerId: string) => {
const key = await mySecureStorage.getKey(providerId);
return key || null;
};
const llmService = new LLMService(myKeyProvider);See Core Concepts for detailed examples including Electron integration.
Logging Configuration
Control logging verbosity via environment variable or service options:
# Environment variable (applies to all services)
export GENAI_LITE_LOG_LEVEL=debug # Options: silent, error, warn, info, debug// Per-service configuration
const llmService = new LLMService(fromEnvironment, {
logLevel: 'debug', // Override env var
logger: customPinoLogger // Inject pino/winston/etc.
});See Logging for custom logger integration and testing patterns.
Example Applications
The library includes two complete demo applications showcasing all features:
- chat-demo - Interactive chat application with all LLM providers, template rendering, and advanced features
- image-gen-demo - Interactive image generation UI with OpenAI and local diffusion support
Both demos are production-ready React + Express applications that serve as reference implementations and testing environments. See Example: Chat Demo and Example: Image Demo for detailed documentation.
Contributing
Contributions are welcome! Please feel free to submit a Pull Request. For major changes, please open an issue first to discuss what you would like to change.
Development
npm install
npm run build
npm testSee Troubleshooting for information about E2E tests and development workflows.
License
This project is licensed under the MIT License - see the LICENSE file for details.
Acknowledgments
Originally developed as part of the Athanor project, genai-lite has been extracted and made standalone to benefit the wider developer community.
