genai-electron
v0.26.0
Published
Electron-first local AI runtime management with Node-safe llama-server launch utilities
Maintainers
Readme
genai-electron
Version: 0.26.0 | Status: Production Ready - Diffusion stuck-job watchdog and wrapper access control
Electron-first library for managing local AI model servers (llama.cpp, stable-diffusion.cpp), with supported Electron-free subpaths for calibration policy metadata and direct llama-server launch. Complements genai-lite for API abstraction.
Features
- ✅ System detection - Auto-detect RAM, CPU, GPU, VRAM capabilities
- ✅ Model management - Download GGUF models with pinned Hugging Face revisions, persisted provenance, progress tracking, and metadata extraction
- ✅ LLM server - Manage llama-server lifecycle with auto-configuration
- ✅ LLM runtime calibration - Find a best-known start-ready configuration within a host-selected time across one or two comparable context profiles, with explicit evidence/completeness and exact caller-supplied diagnostics
- ✅ Image generation - Local image generation via a persistent stable-diffusion.cpp
sd-serverbackend with single/burst VRAM residency - ✅ Multi-component diffusion models - Flux 2, SDXL split components with aggregate checksum validation
- ✅ Resource orchestration - Symmetric: automatic LLM offload/reload when memory constrained, and a resident diffusion backend yields to an LLM start
- ✅ Reliable server lifecycle - Crash auto-restart, hang and stuck-job watchdogs, automatic port selection, occupancy safety, log rotation
- ✅ Advanced launch control - KV-cache quantization, flash-attention control, MoE CPU offload, multi-shard GGUF downloads, image-generation cancellation
- ✅ Binary management - Automatic binary download with GPU variant testing (CUDA→Vulkan→CPU)
- ✅ TypeScript-first - Full type safety, minimal runtime dependencies
Installation
npm install genai-electron
npm install electron@>=25.0.0 # Required only for the package root/manager APIsQuick Start
import { app } from 'electron';
import { systemInfo, modelManager, llamaServer } from 'genai-electron';
app.whenReady().then(async () => {
// Detect capabilities
const caps = await systemInfo.detect();
console.log('RAM:', (caps.memory.total / 1024 ** 3).toFixed(1), 'GB');
// Download model (if needed)
await modelManager.downloadModel({
source: 'url',
url: 'https://huggingface.co/TheBloke/Llama-2-7B-GGUF/resolve/main/llama-2-7b.Q4_K_M.gguf',
name: 'Llama 2 7B',
type: 'llm'
});
// Start server with auto-config (port is optional — defaults to 8080;
// pass port: 'auto' to have the OS pick a free port)
const info = await llamaServer.start({
modelId: 'llama-2-7b'
});
console.log(`Server ready on port ${info.port}`);
});Use with genai-lite for AI interactions (chat, image generation).
Electron-free calibration policy metadata
Plain-Node tests and build scripts can inspect the current LLM calibration policy without loading the Electron-specific package root:
import { LLAMA_CALIBRATION_DEFAULTS } from 'genai-electron/llm-calibration-policy';
console.log(LLAMA_CALIBRATION_DEFAULTS.policyVersion);Use this supported subpath when validating persisted calibration compatibility outside an Electron runtime.
Electron-free llama-server launch
Plain Node ESM applications can launch a caller-provided binary and GGUF without importing the Electron-backed managers:
import { startLlamaServerRunner } from 'genai-electron/llama-server-launch';
const server = await startLlamaServerRunner({
binaryPath: '/opt/llama/bin/llama-server',
model: { path: '/opt/models/model.gguf' },
config: { host: '127.0.0.1', gpuLayers: 40 },
contextSize: 8192,
parallelRequests: 2,
startupTimeoutMs: 120_000,
port: 12_345,
slotsEndpoint: 'disabled',
});
try {
console.log(server.port, server.capacity);
} finally {
await server.stop();
}Electron is an optional peer for these two Node-safe subpaths. The supported loader contract is
native ESM; a CommonJS build that rewrites dynamic import() to require() must preserve native
import or use an ESM bridge.
Documentation
📚 Complete Documentation - Full API reference, guides, and examples
Quick Links:
- Installation & Setup
- System Detection
- Model Management
- LLM Server
- Image Generation
- TypeScript Reference
- Troubleshooting
Example App
See electron-control-panel for a full-featured reference implementation.
Platform Support
- macOS: 11+ (Intel, Apple Silicon with Metal)
- Windows: 10+ (64-bit; CUDA, Vulkan, CPU)
- Linux: Ubuntu 20.04+, Debian 11+, Fedora 35+ (Vulkan, CPU — no CUDA prebuilt upstream; build from source for CUDA)
License
MIT License - see LICENSE for this project and THIRD_PARTY_NOTICES.md for embedded third-party code.
Related Projects
- genai-lite - Lightweight API abstraction for AI providers (cloud and local)
