openvino-genai-node
v2026.4.1
Published
OpenVINO™ GenAI pipelines for using from Node.js environment
Maintainers
Readme
OpenVINO™ GenAI for Node.js
Getting Started • Quick install • AI Scenarios • Documentation • Supported Models
openvino-genai-node brings OpenVINO™ GenAI with the most popular Generative AI model pipelines to Node.js. Run LLMs, VLMs, image and speech generation, and RAG locally with hardware-accelerated inference on top of OpenVINO Runtime.
This library is friendly to PC and laptop execution, and optimized for resource consumption. Prebuilt native addons, the OpenVINO runtime and tokenization are fetched during npm install. You do not need a separate OpenVINO SDK install for typical use.
Key Features and Benefits:
- 📦 Pre-built Generative AI Pipelines: Ready-to-use pipelines for text generation (LLMs), visual language models (VLMs), image generation (Diffusers), speech recognition (Whisper), and speech generation (SpeechT5). See all supported AI scenarios.
- 👣 Minimal Footprint: Smaller binary size and reduced memory footprint compared to other frameworks.
- 📥 Plug-and-play install: Prebuilt native addons and OpenVINO runtime are downloaded on
npm install— no manual OpenVINO installation for typical use. - 🖥️ In-process inference: Run models directly from Node.js in your application process — no separate inference service or access tokens.
- 🚀 Performance Optimization: Hardware-specific optimizations for CPU, GPU, and NPU devices. See optimization techniques in the GenAI docs.
- 👨💻 Programming Language Support: APIs aligned with Python/C++ where possible — async functions and ESM.
- 🗜️ Model Compression: Support for 8-bit and 4-bit weight compression, including embedding layers.
- 🎓 Advanced Inference Capabilities: In-place KV-cache, dynamic quantization, speculative sampling, and more.
- 🎨 Wide Model Compatibility: Support for popular models including Llama, Mistral, Phi, Qwen, Stable Diffusion, Flux, Whisper, and others. Refer to the Supported Models for more details.
Quick install
npm install openvino-genai-nodeRequirements
Node.js ≥ 21. Refer to the supported platforms for more details.
Supported platforms
| OS | x64 | arm64 | | ------ | ------- | --------- | | Windows | ✅ | ❌ | | Linux | ✅ | ✅ | | macOS | ❌ | ✅ |
Prebuilt binaries are downloaded for your OS/arch during npm install. If a platform is unsupported, installation or loading the native addon will fail — build from source per BUILD.md.
Getting started
Model preparation
Models must be converted to OpenVINO IR (for example with Optimum Intel). See the GenAI documentation for model preparation and supported architectures.
Minimal example (LLM)
Import a pipeline from openvino-genai-node, point it at a prepared model directory, then call generate or stream:
import { LLMPipeline } from "openvino-genai-node";
const modelPath = "/path/to/ov/model"; // directory with OpenVINO IR + tokenizer files
const device = "CPU"; // or "GPU", "NPU" where supported
const config = { max_new_tokens: 100 };
const pipe = await LLMPipeline(modelPath, device);
// One-shot generation
const result = await pipe.generate("What is OpenVINO?", config);
console.log(result);
// Streaming tokens to stdout
for await (const chunk of pipe.stream("What is OpenVINO?", config)) {
process.stdout.write(chunk);
}More runnable examples: samples/js.
Supported Generative AI Scenarios
The OpenVINO™ GenAI Node.js package supports the following scenarios (details and model lists are in the linked use-case docs):
- Text generation using Large Language Models (LLMs) - Chat with local Llama, Phi, Qwen and other models
- Visual processing using Vision Language Model (VLMs) - Analyze images/videos with LLaVa, MiniCPM-V and other models
- Image generation using Diffusers - Generate images with Stable Diffusion & Flux models
- Speech recognition using Whisper - Convert speech to text using Whisper models
- Speech generation using SpeechT5 - Convert text to speech using SpeechT5 TTS models
- Semantic search using Text Embedding - Compute embeddings for documents and queries to enable efficient retrieval in RAG workflows
- Text Rerank for Retrieval-Augmented Generation (RAG) - Analyze the relevance and accuracy of documents and queries for your RAG workflows
Contributing
Contributions are welcome in the openvino.genai repository.
- Read the project Contributing guide (PR process, code quality, branching).
- Node.js bindings: build and test locally — BUILD.md (native addon + TypeScript wrapper in
src/js). - JavaScript changes: from
src/js, runnpm run lint,npm run build, andnpm test(tests require Python 3.10+ and model setup described in BUILD.md). - Samples: extend or fix examples under samples/js.
- Open issues or pull requests on GitHub; use the repository PR template and fork-based workflow described in the contributing guide.
FAQ
Can I use OpenVINO GenAI in the browser?
Not with openvino-genai-node — this package targets Node.js (server-side) only. For in-browser inference, use Transformers.js.
Can I use OpenVINO GenAI for Node.js as an ML inference server?
No. openvino-genai-node does not expose REST or gRPC APIs; it runs models in-process inside your Node.js application. For a dedicated model-serving layer with HTTP/gRPC, deploy OpenVINO Model Server (OVMS) and call it from Node (or other clients).
License
The OpenVINO™ GenAI repository is licensed under Apache License Version 2.0. By contributing to the project, you agree to the license and copyright terms therein and release your contribution under these terms.
