@codingmanfocus/alri
v0.1.2
Published
Auto LLM Runtime Installer
Downloads
42
Maintainers
Readme
ALRI
ALRI is the Auto LLM Runtime Installer: a conservative Node.js and TypeScript library and CLI for installing LLM runtimes, resolving models, and running foreground model servers with safe defaults.
The project currently targets:
- llama.cpp — stable, first-class support on macOS, Linux, and Windows.
- SGLang — beta support on Linux with NVIDIA GPUs.
- vLLM — beta support on Linux with NVIDIA GPUs.
ALRI has no runtime npm dependencies or native Node.js addons. It uses Node.js standard-library APIs and OS tools only where runtime management genuinely requires them.
Status
ALRI is under active development. Its npm tarball and external consumer surface are verified before publication. llama.cpp is the primary compatibility target. SGLang and vLLM are intentionally marked beta because their Python, CUDA, driver, and GPU compatibility constraints are more environment-dependent.
Requirements
- Node.js 22 or newer
- npm 10 or newer
- TypeScript 5.5 or newer when consuming the published declarations directly
- For llama.cpp source builds: an exact official
b3143or newer tag, Git, CMake, a supported compiler, and the SDK for the selected backend - For SGLang/vLLM: glibc-based Linux, CPython 3.10–3.13 with
venvandensurepip, a supported NVIDIA GPU, and a sufficiently recent NVIDIA driver
ALRI never installs or modifies system Python, GPU drivers, CUDA, ROCm, Vulkan, or system SDKs. It detects available host capabilities, installs only managed runtime artifacts or isolated Python environments, and reports actionable diagnostics when a required system component is missing.
Library API
Install ALRI as a library:
npm install @codingmanfocus/alriThe package root exposes a deliberately small, typed API. The normal application flow is to ensure a host-appropriate runtime is installed and selected, start a model server, use its OpenAI-compatible client, then stop it explicitly:
import { createAlri } from "@codingmanfocus/alri";
const alri = createAlri();
await alri.runtimes.ensure({
runtime: "llama.cpp",
backend: "auto",
});
const server = await alri.servers.start({
runtime: "llama.cpp",
model: "hf://owner/repository/model.gguf?revision=main",
host: "127.0.0.1",
port: 8080,
ctxSize: 8192,
gpuLayers: "auto",
});
try {
const response = await server.inference.chat({
model: server.servedModelName,
messages: [{ role: "user", content: "Hello" }],
});
console.log(response.choices[0]?.message.content);
} finally {
await server.stop();
}For a prebuilt llama.cpp request, runtimes.ensure() first reuses a selected
matching managed installation unless reinstall: true is explicit;
runtimes.install() bypasses that selected-first shortcut and follows normal
installer resolution, while reinstall: true forces replacement. Source and
Python installers apply their own exact-plan reuse rules. servers.start()
resolves and verifies the GGUF for llama.cpp, waits for runtime-specific
readiness and model identity checks, then returns a handle. server.closed
reports a later natural exit or failure; server.stop() is idempotent and waits
for complete process-tree cleanup.
The built-in inference client currently provides listModels() and bounded,
non-streaming chat() calls. It maps camelCase options to the OpenAI-compatible
wire format and retains runtime-specific fields under raw. For streaming or
other endpoints, use server.apiBaseUrl with another client or native fetch.
An AbortSignal can cancel installs, model resolution, server lifetime, and
inference requests without replacing the caller's cancellation reason.
SGLang and vLLM use model IDs or local model directories directly:
await alri.runtimes.ensure({ runtime: "vllm", version: "0.26.0" });
const server = await alri.servers.start({
runtime: "vllm",
model: "owner/model",
servedModelName: "local-model",
host: "127.0.0.1",
port: 8000,
});
try {
await server.inference.listModels();
} finally {
await server.stop();
}createAlri() accepts optional normalized absolute rootDirectory and
workingDirectory paths plus in-memory githubToken and huggingFaceToken
values. Construction itself does not create managed directories.
Models and diagnostics are also available without routing through the CLI:
const resolved = await alri.models.resolve("./models/model.gguf");
await alri.models.register({
name: "local-model",
reference: resolved.localPath,
});
const registeredModels = await alri.models.list();
const diagnosticReport = await alri.diagnostics.inspect();Connect to an already-running OpenAI-compatible server with either
alri.createInferenceClient() or the equivalent root export:
const client = alri.createInferenceClient({
baseUrl: "http://127.0.0.1:8080",
});
const models = await client.listModels();Build and run
Install the CLI globally from npm:
npm install --global @codingmanfocus/alri
alri --helpnpm ci
npm run build
node dist/src/cli.js --helpThe examples below use alri for readability. During development, substitute
node dist/src/cli.js, or use npm link if you explicitly want a global local
link.
llama.cpp
Install the latest official release artifact for the current host:
alri install llama.cpp
alri install llama.cpp --backend cpu
alri install llama.cpp --version b10250 --backend metalOmitting --backend, or selecting auto, probes the host conservatively.
Apple silicon selects Metal; other supported hosts choose a managed CUDA, ROCm,
or Vulkan candidate only when its runtime or toolchain command succeeds, and
otherwise fall back to CPU. ALRI then verifies the exact official artifact and
runs the installed executable, but a command probe is not a guarantee that
every model and device combination will work. Source builds require a complete
compiler toolchain for an accelerated backend. An explicit backend overrides
automatic selection and remains subject to platform and artifact validation.
The explicit Metal example above is for macOS on Apple silicon.
ALRI resolves exact official asset names, requires the release-provided SHA-256 digest, downloads without credential-bearing redirects, extracts into staging, verifies the installed executables, then activates the installation atomically.
Build an exact official tag from source:
alri install llama.cpp \
--source \
--version b10250 \
--backend cpu \
--jobs 8Source builds verify that the checked-out HEAD is the exact requested tag
commit. ALRI validates the checked-out CMake layout and backend declarations
before configuring, records the detected source-build dialect in provenance,
and fails closed when a tag does not expose the requested capability. Tags
older than b3143 are outside the supported source-build dialect boundary.
Managed serving deliberately has a newer compatibility floor: official
b10250 or later. Older managed builds remain available through alri run,
but ALRI rejects them before server launch because historical server option
names and value arity cannot satisfy the current managed safety contract
reliably.
Supported prebuilt backends depend on the official release matrix and host:
- macOS: CPU and Metal
- Linux: CPU, Vulkan, ROCm, OpenVINO, and SYCL where official artifacts exist
- Windows: CPU, CUDA, Vulkan, OpenVINO, SYCL, HIP, and ARM64 OpenCL where official artifacts exist
Use --source for supported combinations that do not have an official prebuilt
artifact.
GGUF model management
ALRI accepts local GGUF paths, HTTP(S) URLs, compact Hugging Face repository and quantization references, explicit Hugging Face file references, and registered model names:
alri models resolve ./models/model.gguf
alri models resolve https://example.invalid/model.gguf --sha256 <digest>
alri models resolve 'unsloth/Qwen3.5-4B-GGUF@Q8_0'
alri models resolve 'hf://owner/repository/model.gguf?revision=main'
alri models add tiny ./models/tiny.gguf
alri models add managed ./models/model.gguf --copy --sha256 <digest>
alri models add qwen35 'unsloth/Qwen3.5-4B-GGUF@Q8_0'
alri models listThe compact owner/repository@quantization form finds the unique GGUF whose
filename ends in that quantization (case-insensitively). Use an explicit hf://
reference when a repository has multiple matching files. Managed blobs are
content-addressed by SHA-256. Hugging Face downloads use LFS digest and size
metadata when it is available. ALRI validates the GGUF fixed
header and basic structural bounds before registering or serving a file.
Serve a GGUF with the selected llama.cpp installation:
alri serve tiny
alri serve ./models/model.gguf --ctx-size 8192 --gpu-layers auto
alri serve tiny --host 0.0.0.0 --allow-networkThe default bind is 127.0.0.1:8080, and the upstream web UI is disabled.
Arguments after -- are passed directly to llama-server without a shell, but
ALRI blocks overrides of managed model, bind, alias, and UI options.
Installed llama.cpp executables are also available without modifying PATH:
alri run llama.cpp
alri run llama.cpp llama -- --help
alri run llama.cpp quantize -- input.gguf output.gguf Q4_K_MSGLang and vLLM
The beta installers create isolated managed virtual environments and pin the top-level runtime package version:
alri install sglang
alri install vllm
alri install vllm --version 0.26.0 --python /usr/bin/python3Current repository defaults are SGLang 0.5.16 with the cu130 wheel profile
and vLLM 0.26.0 with the cu129 wheel profile. ALRI probes CPython, glibc,
NVIDIA driver compatibility, and GPU compute capability before installation.
Serve a Hugging Face model ID or local model directory directly:
alri serve meta-llama/Llama-3.1-8B-Instruct \
--runtime sglang \
--tensor-parallel-size 2
alri serve /models/Qwen \
--runtime vllm \
--gpu-memory-utilization 0.9 \
--served-model-name qwen-localSGLang defaults to 127.0.0.1:30000; vLLM defaults to
127.0.0.1:8000. Remote repository code is disabled unless
--trust-remote-code is explicitly supplied. Arbitrary Python-runtime
passthrough is intentionally disabled because upstream aliases could otherwise
override ALRI-managed bind policy.
Diagnostics and JSON output
alri doctor
alri runtimes
alri runtimes --json
alri models list --jsonalri doctor reports Node.js, source-build tools, Python, accelerator commands,
and managed installation state. Runtime listing, installation, diagnostics,
model management, server readiness, and managed tool listing support structured
JSON where applicable; raw upstream process output deliberately does not.
Storage and environment
Set ALRI_HOME to use an explicit managed root. Otherwise ALRI uses:
- macOS:
~/Library/Application Support/alri - Linux:
$XDG_DATA_HOME/alrior~/.local/share/alri - Windows:
%LOCALAPPDATA%\alri
Optional credentials:
GITHUB_TOKENfor official GitHub release API requestsHF_TOKENfor Hugging Face metadata, model downloads, and gated model access by managed SGLang/vLLM servers
Tokens are attached only to their intended origins and are not written into
runtime manifests or the model registry. Generic managed subprocesses do not
inherit either token; Python model servers receive only HF_TOKEN when it is
available.
Safety model
- No shell is used for runtime, build, server, or raw tool commands.
- Archive paths, links, CRC values, and extraction boundaries are validated.
- Downloads use bounded redirect handling, timeouts, staging files, and digest checks where an authoritative digest exists.
- Installations and selections use file locks and atomic state replacement.
- Foreground servers hold an ALRI endpoint lease for their complete lifetime; readiness is announced only after the responding server's runtime-specific model identity matches the launch request.
- Servers bind to loopback unless
--allow-networkis explicit. - Ctrl-C and timeout cancellation terminate the complete process tree, escalate
after a grace period, clean temporary files, and preserve conventional exit
statuses (
130for SIGINT and143for SIGTERM). - Reinstallation keeps the previous active runtime until the replacement is fully verified and selected.
Architecture
ALRI keeps policy and side effects separated by responsibility:
src/domaindefines validated state and errors.src/applicationcoordinates one use case at a time.src/apidefines the stable package-facing ports and adapters.src/infrastructureowns filesystem, HTTP, process, lock, and archive I/O.src/runtimecontains runtime-specific installation and command policy.src/cliparses and renders;src/bootstrapis the composition root.
Runtime modules depend on shared contracts rather than on one another. New runtime integrations should preserve this direction and avoid adding an npm runtime dependency when a bounded Node.js standard-library implementation is practical.
Development
npm run check
npm test
npm run verify:package
npm run verifyStable npm releases are published from GitHub Releases. The release tag must be
exactly v followed by the version in package.json, such as v0.1.0.
Pre-releases are deliberately not published by the stable release workflow.
Package verification runs npm pack, installs the produced tarball into an
empty offline consumer project, checks the root import and private deep-import
boundary, and compiles an external strict TypeScript consumer against the
published declarations.
The TypeScript configuration enables strict checking, unchecked-index access checks, exact optional properties, unused-code checks, and fallthrough checks. Keep application services single-purpose, prefer camelCase names, and add a focused regression test for every compatibility or safety boundary.
