@authuser/airllm
v0.2.0
Published
Use AirLLM (layer-by-layer LLM inference for low-memory GPUs) from Node.js, with zero manual Python setup.
Maintainers
Readme
@authuser/airllm
Run AirLLM — layer-by-layer inference that lets large LLMs run on low-VRAM hardware — from Node.js and TypeScript, with zero manual Python setup.
Requirements
- Node.js >= 18. The package is ESM-only (
"type": "module") and ships its own TypeScript types. - Windows, macOS, or Linux. No browser/edge support — it runs a local Python process.
- No Python installation needed. The first call to
AirLLM.load()provisions an isolated Python runtime in your user data directory (it never touches your system Python). This needs network access and can take several minutes and a few GB the first time; later loads reuse what was already installed. - NVIDIA GPU (CUDA) verified on real hardware; CPU is fully supported; Apple Silicon (MLX) is implemented but not yet verified on real hardware.
Installation
npm install @authuser/airllmQuick start
import { AirLLM } from '@authuser/airllm';
const model = await AirLLM.load('some-org/some-sharded-model', {
device: 'cpu', // 'cuda' | 'cpu' | 'mlx' — required, load() does not auto-detect
compression: '4bit', // optional: 4-bit/8-bit quantization
});
try {
for await (const token of model.generate('Hello, how are you?', { maxTokens: 200 })) {
process.stdout.write(token);
}
} finally {
await model.unload();
}Each AirLLM instance holds a single active model. To switch models, unload() the current
one and AirLLM.load() again.
Streaming and cancellation
const controller = new AbortController();
for await (const token of model.generate('...', { signal: controller.signal })) {
// controller.abort() at any point stops the stream
}API
AirLLM.load(modelId: string, options?: LoadOptions): Promise<AirLLM>
| Option | Type | Default | Notes |
|---|---|---|---|
| device | 'cuda' \| 'cpu' \| 'mlx' | AirLLM default: cuda:0 | No auto-detection — pass it explicitly |
| compression | '4bit' \| '8bit' | no quantization | |
| hfToken | string | — | Needed for gated Hugging Face repos |
| prefetching | boolean | true | |
| maxSeqLen | number | 512 | Prompts are truncated to this length |
| profilingMode | boolean | false | |
model.generate(prompt: string, options?: GenerateOptions): AsyncGenerator<string>
| Option | Type | Notes |
|---|---|---|
| maxTokens | number | Maximum number of new tokens |
| temperature | number | Enables sampling; without it generation is deterministic |
| signal | AbortSignal | Cancels the in-progress generation |
model.unload(): Promise<void>
Releases the model and stops the underlying Python process. The instance can't be reused
afterwards — call AirLLM.load() again.
model.getMetrics(): Metrics
| Field | Type | Notes |
|---|---|---|
| tokensPerSecond | number | |
| memoryUsedMB | number \| null | null outside of CUDA (CPU/MLX) |
| activeRequests | number | |
| uptimeMs | number | Since the instance was created |
Events
AirLLM extends EventEmitter:
'log'— raw log lines from the underlying process.'model:download:progress'—{ status: 'downloading' | 'ready' }.'metrics'— aMetricssnapshot emitted periodically during generation.'crash'— the underlying process died unexpectedly; the instance becomes unusable and any latergenerate()/unload()call throwsGenerationError. CallAirLLM.load()again.
Errors
All error classes extend Error and are exported from the package root:
| Class | Extends | Thrown when |
|---|---|---|
| ModelLoadError | Error | load() fails |
| OutOfMemoryError | ModelLoadError | Out of memory while loading |
| AuthenticationError | ModelLoadError | Missing/invalid Hugging Face token for a gated repo |
| GenerationError | Error | generate()/unload() fails, or the instance is no longer usable |
| SidecarCrashError | GenerationError | The underlying process died mid-operation |
Examples
The examples/ directory has runnable, non-mocked demos:
basic-generate.mjs— load, stream, cancel, and unload on CPU.gpu-generate.mjs— the same flow on an NVIDIA GPU, reporting realmemoryUsedMBusage.
Roadmap
- Concurrent loading of multiple models
- Per-model-family classes on top of the current generic API
- OpenTelemetry integration
- Injectable structured JSON logging
- Automatic process restart after a crash
- End-to-end verification on Apple Silicon / MLX
License
MIT — see LICENSE.
