@veloxquant/sdk
v0.7.1
Published
JavaScript/TypeScript SDK for VeloxQuant-MLX: hardware-aware memory optimization and OpenAI-compatible local inference on Apple Silicon.
Maintainers
Readme
@veloxquant/sdk
The JavaScript/TypeScript SDK for VeloxQuant-MLX. Hardware-aware KV-cache memory optimization and OpenAI-compatible local inference on Apple Silicon — without writing Python.
Overview
@veloxquant/sdk is a typed bridge, not a reimplementation: it shells out to
the veloxquant CLI and talks HTTP to a veloxquant serve process that it
manages for you, so quantization and hardware detection stay the
responsibility of the underlying veloxquant-mlx engine.
Contents
- Overview
- Requirements
- Install
- Quick start
- Streaming
- Autopilot
- Memory calculator
- Compression methods
- Local model management
- OpenAI compatibility
- Vercel AI SDK
- LangChain.js
- LlamaIndex.TS
- Tool-calling agent
- Hardware-aware memory optimization
- Persistent model sessions
- Multi-turn conversations
- Structured output / JSON mode
- Benchmarking
- CLI
- Known limitations
- Development
- License
Requirements
- macOS on Apple Silicon (M1 or later) — MLX has no other backend.
- Python 3.10+ with
veloxquant-mlxinstalled (pip install veloxquant-mlx). - Node.js 18.17+.
Install
npm install @veloxquant/sdkQuick start
import { VeloxQuant } from "@veloxquant/sdk";
const vq = new VeloxQuant();
const response = await vq.chat({
model: "mlx-community/Qwen3-8B-4bit",
messages: [{ role: "user", content: "Explain quantum computing simply." }],
});
console.log(response.text);Streaming
for await (const chunk of vq.stream({
model: "mlx-community/Qwen3-4B-4bit",
prompt: "Explain transformers",
})) {
if (!chunk.done) process.stdout.write(chunk.text);
}Autopilot
autopilot() is the zero-config entry point for "I don't want to think about
memory." Give it a model and a rough model-size class; it gets a
hardware-aware recommendation for this machine, refuses to start a server for
a configuration that clearly won't fit, and loads the model with a method
veloxquant serve can actually run:
import { autopilot } from "@veloxquant/sdk";
const ai = await autopilot({ model: "mlx-community/Llama-3.2-1B-Instruct-4bit", modelClass: "1B" });
const answer = await ai.chat({ prompt: "Build a REST API in TypeScript" });
await ai.stop();If the recommended configuration is unlikely to fit in available memory,
autopilot() throws an AutopilotFitError describing why instead of starting
a server that would likely crash or thrash:
try {
await autopilot({ model: "mlx-community/Qwen3-32B-4bit", modelClass: "32B" });
} catch (err) {
if (err instanceof AutopilotFitError) {
console.error(err.message); // lists the specific won't-fit warning(s)
}
}Pass { force: true } to start the server anyway. Note: autopilot() does
not pick a model for you — veloxquant-mlx has no LLM catalog to select
from, only a KV-cache compression config recommender — model must always be
a model you name yourself.
Memory calculator
vq.memory.estimate() answers "will this fit?" for a given attention shape —
fp16 KV-cache size, the compression method veloxquant-mlx would pick, and a
directional estimate of what compression saves:
import { VeloxQuant, formatBytes } from "@veloxquant/sdk";
const vq = new VeloxQuant();
const estimate = await vq.memory.estimate({ seqLen: 32768, headDim: 128, nLayers: 32 });
console.log(`fp16 KV-cache: ${formatBytes(estimate.fp16KvCacheBytes)}`);
console.log(`Recommended method: ${estimate.recommendedMethod}`);
console.log(`Estimated compressed: ${formatBytes(estimate.estimatedCompressedBytes)}`);
console.log(`Estimated saved: ${formatBytes(estimate.memorySavedBytes)}`);fp16 KV-cache: 512 MB
Recommended method: kvquant
Estimated compressed: 96 MB
Estimated saved: 416 MB(Real output from examples/memory-estimate.ts --seq-len 32768 --head-dim 128 --n-layers 32
against veloxquant-mlx 0.71.1 — the method and savings above depend on the
installed package version and the workload shape you pass in.)
A standalone, copy-pasteable version of this lives at
examples/memory-estimate.ts:
npx tsx examples/memory-estimate.ts --seq-len 32768 --head-dim 128 --n-layers 32The same caveat from Known limitations applies here: these byte counts are accounting-only (fidelity estimates), not measured runtime memory freed.
Compression methods
vq.models.list() — despite the name, this lists veloxquant-mlx's KV-cache
compression methods, not LLMs (the package has no model catalog; see the
Overview). Use it to discover what's available and which methods
veloxquant serve can actually run:
const methods = await vq.models.list({ servableOnly: true });
console.log(`Default serve method: ${methods.defaultServeMethod}`);
for (const m of methods.methods.slice(0, 5)) {
console.log(`${m.name.padEnd(16)} ${m.family.padEnd(12)} ${m.serveTierLabel}`);
}Default serve method: turboquant_rvq
a2ats hybrid available
adakv quantization available
age_tiered quantization available (no prompt-cache trimming)
amc eviction available (no prompt-cache trimming)
anchorkv hybrid available (no prompt-cache trimming)(Real output from veloxquant-mlx 0.71.1 — 43 methods total as of this
version.) Filter by family with vq.models.list({ family: "quantization" });
valid families are "quantization", "eviction", and "hybrid".
Every method's byte-count reporting is accounting-only — see
Known limitations — which is also why
methods.accountingNote is included directly in the result rather than left
for callers to discover separately.
Local model management
vq.models.local() lists model weights already downloaded to this machine's
Hugging Face cache — distinct from vq.models.list() above, which lists
compression methods, not model weights.
const local = await vq.models.local();
// [{ id: "mlx-community/Qwen3-4B-4bit", sizeBytes: 15947227, lastUsedAt: Date }, ...]Backed by huggingface_hub's own cache scanner (scan_cache_dir()), which
already resolves HF_HOME/HF_HUB_CACHE and correctly deduplicates the
content-addressed blob layout — not a hand-rolled du over the cache
directory.
vq.models.pull() and vq.models.delete() manage the cache directly —
downloading or evicting a model's weights without starting a
veloxquant serve process:
const pulled = await vq.models.pull("mlx-community/Qwen3-4B-4bit");
// { id: "mlx-community/Qwen3-4B-4bit", sizeBytes: 2278969756 }
const deleted = await vq.models.delete("mlx-community/Qwen3-4B-4bit");
// { id: "mlx-community/Qwen3-4B-4bit", freedBytes: 2278969756 }Both are built on the same huggingface_hub APIs the cache scanner above
already relies on: pull() calls snapshot_download(), and delete() uses
scan_cache_dir() → delete_revisions() → execute() rather than an
rm -rf on a resolved path — the cache's blob layout is content-addressed
and shared across revisions/repos via symlinks, so a naive recursive delete
risks corrupting a different cached model.
vq.models.pull() does not default to VeloxQuantOptions.timeoutMs's
usual 30s CLI timeout — model downloads can take many minutes, so pass
{ timeoutMs } explicitly if you want a bound; by default the call waits as
long as it takes. There's no progress callback in this version:
snapshot_download()'s tqdm-based progress doesn't cross the subprocess
boundary cleanly, so a pull() call is silent until it resolves or throws.
Both are exposed on the CLI too:
npx veloxquant models pull mlx-community/Qwen3-4B-4bit
npx veloxquant models delete mlx-community/Qwen3-4B-4bitHardware-aware memory optimization
const info = await vq.system.info();
// { chip: "Apple M4", unifiedMemoryBytes: ..., veloxquantVersion: "0.71.1", ... }
const estimate = await vq.memory.estimate({ seqLen: 32768, headDim: 128, nLayers: 32 });
// { fp16KvCacheBytes, recommendedMethod, estimatedCompressedBytes, memorySavedBytes, reason }
const picked = await vq.optimize({ profile: "maximum-context", seqLen: 32768 });
// picks a method/bit-width biased toward the long-context band
const recommendation = await vq.recommendModel({ modelClass: "7B", goal: "max_context" });
// chip/ramGb auto-detected from this machine when omitted — pass them explicitly to overridePersistent model sessions
Loading a model with automatic optimization keeps the server process alive across multiple turns instead of spinning one up per call:
const model = await vq.load({ model: "mlx-community/Qwen3-8B-4bit", optimize: "auto" });
const r1 = await model.chat({ prompt: "Hi" });
const r2 = await model.chat({ prompt: "Now summarize that" });
await model.stop();Cancel an in-flight completion or stream without stopping the loaded model by
passing an AbortSignal. The same option is available on agent.run():
const controller = new AbortController();
const stream = model.stream({ prompt: "Write a long story", signal: controller.signal });
for await (const chunk of stream) {
process.stdout.write(chunk.text);
if (shouldStop()) controller.abort();
}
// The model process is still available for another request.
await model.stop();Multi-turn conversations
model.conversation() bookkeeps chat history automatically so you don't have
to hand-roll a ChatMessage[] and re-append it after every turn:
const model = await vq.load({ model: "mlx-community/Qwen3-4B-4bit", optimize: "auto" });
const convo = model.conversation({ system: "You are a terse assistant." });
const r1 = await convo.send("What's the capital of France?");
console.log(r1.text); // "Paris."
const r2 = await convo.send("And its population?");
console.log(r2.text); // has full history context
console.log(convo.messages); // read-only view of accumulated ChatMessage[]
convo.reset(); // clears history, keeps the system promptIf you pass tools to send() and the model calls one, you're responsible
for executing it and feeding the result back as a tool-role message
yourself — Conversation only handles history bookkeeping, not a second
agent loop (see Tool-calling agent for that).
Structured output / JSON mode
Ask the model for JSON back with responseFormat:
const result = await model.chat({
messages: [{ role: "user", content: "Extract the name and age from: John is 30." }],
responseFormat: {
type: "json_schema",
schema: {
type: "object",
properties: { name: { type: "string" }, age: { type: "number" } },
required: ["name", "age"],
},
},
});
console.log(result.json); // { name: "John", age: 30 } — already JSON.parse()'d{ type: "json_object" } works the same way without a schema, for "just give
me valid JSON, any shape."
This is not a hard guarantee. The mlx_lm server this SDK wraps has no
native support for OpenAI's response_format field — it's not
grammar-constrained decoding under the hood. Under the hood this is a
best-effort prompt-injection fallback: formatting instructions (and the
schema, for json_schema) are appended to the outgoing messages, and the
model's text response is parsed as JSON afterward (markdown code fences are
stripped first, since models sometimes wrap JSON in them despite being told
not to). result.json is null when responseFormat wasn't requested; when
it was requested and the model's output fails to parse, chat() throws a
descriptive error instead of silently returning null (which would be
indistinguishable from "not requested"). responseFormat is still sent on
the wire request too, in case a future server version starts honoring it
natively — it's an inert no-op today either way. See issue
#14 for the
investigation into mlx_lm/server.py behind this.
model.conversation().send() accepts the same responseFormat option.
Benchmarking
vq.benchmark() measures tokens/sec, time-to-first-token, and resident
memory (RSS) for a model on this machine — real numbers from your own
hardware, not a README claim:
import { VeloxQuant } from "@veloxquant/sdk";
const vq = new VeloxQuant();
const result = await vq.benchmark({ model: "mlx-community/Llama-3.2-1B-Instruct-4bit" });
console.log(result.toMarkdown());VeloxQuant Benchmark
Model: mlx-community/Llama-3.2-1B-Instruct-4bit
Machine: Apple M4
RAM: 24GB
Tokens/sec: 124.8
TTFT: 237ms
turboquant_rvq resident memory: 1055MB
kivi resident memory: 1052MB
Resident memory reduced: 0%(Real output from a local run — requires real Apple Silicon hardware and the model already downloaded; there's no way to run this in CI.)
Tokens/sec and TTFT are measured at the SDK boundary from stream() chunk
timestamps. The resident-memory numbers are genuine measured RSS of the
veloxquant serve subprocess (via ps) — not the accounting-only byte
counts vq.memory.estimate() reports (see
Known limitations). Because of that, compression is
not guaranteed to reduce measured RSS: toMarkdown() reports an increase
honestly rather than hiding it, since idle RSS mostly reflects model
weights, and KV-cache compression's effect is small or even negative for a
short prompt against a small model — measured directly against
Llama-3.2-1B-Instruct-4bit above.
OpenAI compatibility
A server started with vq.load() already speaks the OpenAI chat-completions
API — this SDK's own chat()/stream() post directly to
${baseUrl}/v1/chat/completions. That means the real
openai npm package works against it
directly, with no wrapper class:
import OpenAI from "openai";
import { VeloxQuant } from "@veloxquant/sdk";
const vq = new VeloxQuant();
const model = await vq.load({ model: "mlx-community/Llama-3.2-1B-Instruct-4bit", optimize: "auto" });
const client = new OpenAI({ baseURL: `${model.baseUrl}/v1`, apiKey: "local" });
const response = await client.chat.completions.create({
model: model.model, // the served model id — NOT model.method (that's the compression method name)
messages: [{ role: "user", content: "Hello" }],
});Two things to get right, both verified against a real running server:
baseURLmust include/v1— theopenaiclient appends/chat/completionsitself.modelin the request must be the actual served model id —model.model, notmodel.method(that's the KV-cache compression method name, e.g."kivi"— passing it asmodelis rejected with a 404, since the server is pinned to one model per process and treatsmodelas a routing check, not a free label).
A full runnable version (including streaming) is at
examples/openai-client.ts. openai is not a
dependency of this package — install it yourself if you want to use it this
way.
Vercel AI SDK
@veloxquant/sdk/ai-sdk wraps an already-loaded model as a Vercel AI SDK
LanguageModelV4, for use with generateText()/streamText():
import { generateText, streamText } from "ai";
import { veloxquant } from "@veloxquant/sdk/ai-sdk";
import { VeloxQuant } from "@veloxquant/sdk";
const vq = new VeloxQuant();
const model = await vq.load({ model: "mlx-community/Llama-3.2-1B-Instruct-4bit", optimize: "auto" });
const result = await generateText({ model: veloxquant(model), prompt: "Hello!" });
console.log(result.text);
await model.stop();veloxquant() takes an already-running VeloxQuantModel, not a bare model
name — createOpenAICompatible() (the AI SDK helper this is built on) is
synchronous and needs a baseUrl up front, but getting one requires an
async vq.load() first, and an AI SDK provider has no disposal hook where
this function could safely call model.stop() for you. Loading (and
stopping) the model stays the caller's responsibility, same as everywhere
else in this SDK.
ai and @ai-sdk/openai-compatible are optional peer dependencies — install
them yourself to use this subpath. A full runnable version (including
streaming) is at examples/ai-sdk.ts.
LangChain.js
@veloxquant/sdk/langchain wraps an already-loaded model as a LangChain.js
BaseChatModel, for use with invoke(), LCEL chains, and agents:
import { VeloxQuant } from "@veloxquant/sdk";
import { VeloxQuantChatModel } from "@veloxquant/sdk/langchain";
const vq = new VeloxQuant();
const model = await vq.load({ model: "mlx-community/Qwen3-4B-4bit", optimize: "auto" });
const chatModel = new VeloxQuantChatModel(model);
const result = await chatModel.invoke("Explain quantum computing simply.");
console.log(result.content);
await model.stop();Works in LCEL chains the same way any other BaseChatModel does:
import { ChatPromptTemplate } from "@langchain/core/prompts";
import { StringOutputParser } from "@langchain/core/output_parsers";
const prompt = ChatPromptTemplate.fromMessages([
["system", "You are a terse assistant."],
["human", "{question}"],
]);
const chain = prompt.pipe(chatModel).pipe(new StringOutputParser());
console.log(await chain.invoke({ question: "What's the capital of France?" }));Same lifecycle rule as @veloxquant/sdk/ai-sdk: VeloxQuantChatModel takes
an already-running VeloxQuantModel, not a bare model name, since loading is
async and this SDK has no disposal hook to call model.stop() on your
behalf — that stays the caller's responsibility.
Streaming works. .stream() and LCEL chains stream real token deltas
via _streamResponseChunks(), backed by the same SSE stream
VeloxQuantModel.stream() uses:
const stream = await chatModel.stream("Explain quantum computing simply.");
for await (const chunk of stream) {
process.stdout.write(chunk.content as string);
}Tool-call deltas stream too — verified against mlx_lm's server, each
streamed tool call arrives as one complete { id, name, arguments } object
per SSE event rather than an incrementally-assembled fragment, so
tool_call_chunks on each AIMessageChunk are always immediately valid on
their own, not partial JSON needing further concatenation.
One asymmetry vs. invoke(): mlx_lm's server does not send a usage
object on any streamed SSE event (only on the non-streaming response), so
token-usage metadata is unavailable when streaming — llmOutput carries no
tokenUsage on the streaming path, unlike _generate().
@langchain/core is an optional peer dependency — install it yourself to use
this subpath. A full runnable version (direct invoke() and an LCEL chain,
verified against a real running server) is at
examples/langchain.ts.
LlamaIndex.TS
@veloxquant/sdk/llamaindex wraps an already-loaded model as a
LlamaIndex.TS LLM, usable anywhere LlamaIndex
expects one — query engines, agents, Settings.llm:
import { VeloxQuant } from "@veloxquant/sdk";
import { VeloxQuantLLM } from "@veloxquant/sdk/llamaindex";
const vq = new VeloxQuant();
const model = await vq.load({ model: "mlx-community/Qwen3-4B-4bit", optimize: "auto" });
const llm = new VeloxQuantLLM(model);
const response = await llm.chat({ messages: [{ role: "user", content: "Hello!" }] });
console.log(response.message.content);
await model.stop();Streaming works from day one in this adapter (unlike the LangChain.js adapter's original gap):
const stream = await llm.chat({ messages: [{ role: "user", content: "Hi" }], stream: true });
for await (const chunk of stream) process.stdout.write(chunk.delta);complete() (LlamaIndex's plain-prompt API) works too, built on the same
chat() under the hood via LlamaIndex's own BaseLLM:
const completion = await llm.complete({ prompt: "The capital of France is" });
console.log(completion.text);Same lifecycle rule as the other two adapters: VeloxQuantLLM takes an
already-running VeloxQuantModel, not a bare model name — loading is async
and this SDK has no disposal hook to call model.stop() on your behalf.
contextWindow cannot be verified against the served model —
veloxquant-mlx has no LLM catalog (see Overview), so there's
no source of truth for a model's actual context length here. It defaults to
LlamaIndex's own DEFAULT_CONTEXT_WINDOW (3900) unless you pass
{ contextWindow } to the constructor — override it if the served model's
real window matters to your use case (e.g. LlamaIndex sizing its own
chunking/retrieval logic off this value).
Text-only, no tool-calling in this adapter. LlamaIndex's tool metadata
schema doesn't line up with the OpenAI tools shape this SDK's Agent and
wire format use elsewhere, and reconciling the two needs a deliberate design
pass — out of scope for this version. A multimodal message's non-text
content parts (image/audio) are dropped rather than thrown on, so a mixed
text+image message doesn't crash a text-only request outright — but no
non-text content is ever sent to the model.
llamaindex is an optional peer dependency — install it yourself to use
this subpath. A full runnable version (chat, streaming chat, and complete)
is at examples/llamaindex.ts.
Tool-calling agent
vq.agent() loads a model and gives you a single-turn tool-calling loop:
register tools, call run(), and the model calls tools, gets their results
fed back, and produces a final answer.
const agent = await vq.agent({ model: "mlx-community/Qwen3-4B-4bit", optimize: "auto" });
agent.tool({
name: "get_weather",
description: "Get the current weather for a location",
parameters: {
type: "object",
properties: { location: { type: "string", description: "City name" } },
required: ["location"],
},
execute: async ({ location }) => ({ location, tempC: 22, condition: "sunny" }),
});
const result = await agent.run("What's the weather in Tokyo? Do I need an umbrella?");
console.log(result.text);
// "The current weather in Tokyo is 22°C with sunny conditions. ... you likely don't need an umbrella."
await agent.stop();This is built on real, model-native tool-calling — not prompt-engineered
parsing in this SDK. veloxquant serve wraps mlx_lm.server, which checks
the model's own tokenizer for has_tool_calling and parses tool_calls out
of generation using the tokenizer's own chat template and tool-call tokens.
Tool definitions and results use the exact OpenAI tools/tool_calls wire
shape (verified against a real running server with
mlx-community/Qwen3-4B-4bit) rather than a bespoke schema, so the same
tool definitions work whether the request goes through agent.tool(),
vq.chat({ tools }) directly, or the raw OpenAI client from
OpenAI compatibility.
Not every model supports tool-calling — it depends on the model's tokenizer having a tool-calling chat template. If a model doesn't call your tool as expected, that's the first thing to check.
Scope: single-turn tool calling only (call tools → feed results back →
repeat until the model stops calling tools or maxSteps, default 8, is
reached). No multi-step planning beyond that loop — a natural follow-up but a
separate, larger scope.
A full runnable version is at examples/agent.ts.
MCP tool sources
agent.useMcpServer() connects to an MCP
server and registers its tools alongside anything registered with
agent.tool() — both share one dispatch loop in run():
const agent = await vq.agent({ model: "mlx-community/Qwen3-4B-4bit", optimize: "auto" });
await agent.useMcpServer({
name: "filesystem",
transport: "stdio",
command: "npx",
args: ["-y", "@modelcontextprotocol/server-filesystem", "/path/to/allowed/dir"],
});
const result = await agent.run("List the files in the allowed directory.");
console.log(result.text);
await agent.stop(); // also closes any MCP servers useMcpServer() connected itselfOr hand it an already-connected Client from @modelcontextprotocol/sdk
directly ({ name, client }) — in that case the connection's lifecycle stays
with whoever created it, and agent.stop() will not close it.
Can be called more than once, including after the agent has started
running, so a long-lived agent can pick up more tools mid-session.
Registering a tool whose name collides with an already-registered one
(manual or from another MCP server) throws — same behavior as calling
agent.tool() twice with the same name.
@modelcontextprotocol/sdk is an optional peer dependency, loaded lazily
only when useMcpServer() is actually called — the rest of this SDK works
without it installed. Scope: MCP tools only, no resources or prompts
primitives. A tool result with unsupported content (image/audio/resource)
throws a clear error rather than silently dropping it, since surfacing
non-text content to a text-only chat model needs a deliberate design
decision this SDK hasn't made yet.
CLI
npx veloxquant doctor # checks Apple Silicon + veloxquant-mlx install
npx veloxquant recommend # shows detected hardware + servable methods
npx veloxquant recommend --model-class 7B --goal everyday # full recommendation (chip/RAM auto-detected)
npx veloxquant analyze --seq-len 32768 --head-dim 128 --n-layers 32
npx veloxquant serve --model mlx-community/Qwen3-4B-4bit # persistent local servervq serve starts a long-running local server from the terminal — useful for
pointing a separate frontend, curl, or another language's OpenAI client at
it without writing a Node script:
npx veloxquant serve --model mlx-community/Qwen3-4B-4bit --optimizeveloxquant serve ready
baseUrl: http://127.0.0.1:53021
model: mlx-community/Qwen3-4B-4bit
method: turboquant_rvq
pid: 61481
Press Ctrl+C to stop.(Real output from a local run.) It reuses the same startServer() the SDK's
vq.load() calls internally — same VELOXQUANT_READY handshake, same
ServeHandle shutdown path — so behavior is identical whether the server is
started from JS or from this CLI command. Ctrl+C (or SIGTERM) stops the
underlying veloxquant serve subprocess cleanly rather than leaving it
orphaned.
--model <id> Model to serve (required)
--method <name> KV-cache compression method (default: turboquant_rvq)
--bits <n> Bit width override
--port <n> Port to listen on (default: a free port)
--host <addr> Host to bind (default: 127.0.0.1)
--optimize Auto-tune method/bits/knobs for detected hardware
--json Print the ready state as JSON instead of formatted textEvery command accepts --json for scripting/CI use — valid, parseable JSON on
stdout with nothing else mixed in:
npx veloxquant doctor --json{
"ready": true,
"platform": { "ok": true, "value": "darwin" },
"appleSilicon": { "ok": true, "chip": "Apple M4" },
"python": { "ok": true, "interpreter": "python3" },
"veloxquantMlx": { "ok": true, "version": "0.70.0" }
}(Real output, verified against a local machine.)
Known limitations
recommend vs. auto-config stay separate. veloxquant-mlx's
recommend CLI (chip/RAM/model-class picker) and auto-config
(workload/hardware → method picker) are two independent selectors in the
underlying package — this SDK exposes both (vq.recommendModel() and
vq.memory / vq.optimize()) rather than merging them, since the package
itself doesn't merge them either.
Compression byte counts are accounting-only. Caches store dequantized
fp16 tensors today, so byte counters measure compression fidelity, not
runtime memory actually freed. memory.estimate() surfaces this as
estimatedCompressedBytes / memorySavedBytes — treat them as directional,
not a guarantee, until the package changes this (see veloxquant-mlx issue
#27).
recommendModel() inputs are constrained to what's installed.
chip/ramGb are validated against whatever the installed veloxquant
recommend CLI actually accepts (SUPPORTED_CHIPS, SUPPORTED_RAM_GB
exported from this package) — M1-M4 and RAM up to 128GB as of
veloxquant-mlx 0.71.1. Passing values outside that set fails with the CLI's
own argparse error rather than being silently widened.
Development
npm install
npm run build # tsup -> dist/
npm test # node --test against test/unit
npm run typecheck
npm run lintReleasing
Releases and release notes are automated with Changesets. Any PR with a user-facing change (new feature, bug fix, breaking change) should include a changeset:
npx changesetThis prompts for a bump type (patch/minor/major) and a summary, and writes a .changeset/*.md file — commit
it with your PR. CI fails PRs that touch src/ without one (use npx changeset add --empty if a change
genuinely needs no release notes).
On merge to master, a GitHub Action opens or updates a "Version Packages" PR that bumps package.json and
writes CHANGELOG.md from pending changesets. Merging that PR builds, publishes to npm, tags the release, and
creates a GitHub Release with the generated notes — no manual changelog writing or tagging required.
License
MIT
