npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@veloxquant/sdk

v0.7.1

Published

JavaScript/TypeScript SDK for VeloxQuant-MLX: hardware-aware memory optimization and OpenAI-compatible local inference on Apple Silicon.

Readme

@veloxquant/sdk

The JavaScript/TypeScript SDK for VeloxQuant-MLX. Hardware-aware KV-cache memory optimization and OpenAI-compatible local inference on Apple Silicon — without writing Python.

Overview

@veloxquant/sdk is a typed bridge, not a reimplementation: it shells out to the veloxquant CLI and talks HTTP to a veloxquant serve process that it manages for you, so quantization and hardware detection stay the responsibility of the underlying veloxquant-mlx engine.

Contents

Requirements

  • macOS on Apple Silicon (M1 or later) — MLX has no other backend.
  • Python 3.10+ with veloxquant-mlx installed (pip install veloxquant-mlx).
  • Node.js 18.17+.

Install

npm install @veloxquant/sdk

Quick start

import { VeloxQuant } from "@veloxquant/sdk";

const vq = new VeloxQuant();

const response = await vq.chat({
  model: "mlx-community/Qwen3-8B-4bit",
  messages: [{ role: "user", content: "Explain quantum computing simply." }],
});

console.log(response.text);

Streaming

for await (const chunk of vq.stream({
  model: "mlx-community/Qwen3-4B-4bit",
  prompt: "Explain transformers",
})) {
  if (!chunk.done) process.stdout.write(chunk.text);
}

Autopilot

autopilot() is the zero-config entry point for "I don't want to think about memory." Give it a model and a rough model-size class; it gets a hardware-aware recommendation for this machine, refuses to start a server for a configuration that clearly won't fit, and loads the model with a method veloxquant serve can actually run:

import { autopilot } from "@veloxquant/sdk";

const ai = await autopilot({ model: "mlx-community/Llama-3.2-1B-Instruct-4bit", modelClass: "1B" });
const answer = await ai.chat({ prompt: "Build a REST API in TypeScript" });
await ai.stop();

If the recommended configuration is unlikely to fit in available memory, autopilot() throws an AutopilotFitError describing why instead of starting a server that would likely crash or thrash:

try {
  await autopilot({ model: "mlx-community/Qwen3-32B-4bit", modelClass: "32B" });
} catch (err) {
  if (err instanceof AutopilotFitError) {
    console.error(err.message); // lists the specific won't-fit warning(s)
  }
}

Pass { force: true } to start the server anyway. Note: autopilot() does not pick a model for you — veloxquant-mlx has no LLM catalog to select from, only a KV-cache compression config recommender — model must always be a model you name yourself.

Memory calculator

vq.memory.estimate() answers "will this fit?" for a given attention shape — fp16 KV-cache size, the compression method veloxquant-mlx would pick, and a directional estimate of what compression saves:

import { VeloxQuant, formatBytes } from "@veloxquant/sdk";

const vq = new VeloxQuant();
const estimate = await vq.memory.estimate({ seqLen: 32768, headDim: 128, nLayers: 32 });

console.log(`fp16 KV-cache:        ${formatBytes(estimate.fp16KvCacheBytes)}`);
console.log(`Recommended method:   ${estimate.recommendedMethod}`);
console.log(`Estimated compressed: ${formatBytes(estimate.estimatedCompressedBytes)}`);
console.log(`Estimated saved:      ${formatBytes(estimate.memorySavedBytes)}`);
fp16 KV-cache:        512 MB
Recommended method:   kvquant
Estimated compressed: 96 MB
Estimated saved:      416 MB

(Real output from examples/memory-estimate.ts --seq-len 32768 --head-dim 128 --n-layers 32 against veloxquant-mlx 0.71.1 — the method and savings above depend on the installed package version and the workload shape you pass in.)

A standalone, copy-pasteable version of this lives at examples/memory-estimate.ts:

npx tsx examples/memory-estimate.ts --seq-len 32768 --head-dim 128 --n-layers 32

The same caveat from Known limitations applies here: these byte counts are accounting-only (fidelity estimates), not measured runtime memory freed.

Compression methods

vq.models.list() — despite the name, this lists veloxquant-mlx's KV-cache compression methods, not LLMs (the package has no model catalog; see the Overview). Use it to discover what's available and which methods veloxquant serve can actually run:

const methods = await vq.models.list({ servableOnly: true });
console.log(`Default serve method: ${methods.defaultServeMethod}`);
for (const m of methods.methods.slice(0, 5)) {
  console.log(`${m.name.padEnd(16)} ${m.family.padEnd(12)} ${m.serveTierLabel}`);
}
Default serve method: turboquant_rvq

a2ats            hybrid       available
adakv            quantization available
age_tiered       quantization available (no prompt-cache trimming)
amc              eviction     available (no prompt-cache trimming)
anchorkv         hybrid       available (no prompt-cache trimming)

(Real output from veloxquant-mlx 0.71.1 — 43 methods total as of this version.) Filter by family with vq.models.list({ family: "quantization" }); valid families are "quantization", "eviction", and "hybrid".

Every method's byte-count reporting is accounting-only — see Known limitations — which is also why methods.accountingNote is included directly in the result rather than left for callers to discover separately.

Local model management

vq.models.local() lists model weights already downloaded to this machine's Hugging Face cache — distinct from vq.models.list() above, which lists compression methods, not model weights.

const local = await vq.models.local();
// [{ id: "mlx-community/Qwen3-4B-4bit", sizeBytes: 15947227, lastUsedAt: Date }, ...]

Backed by huggingface_hub's own cache scanner (scan_cache_dir()), which already resolves HF_HOME/HF_HUB_CACHE and correctly deduplicates the content-addressed blob layout — not a hand-rolled du over the cache directory.

vq.models.pull() and vq.models.delete() manage the cache directly — downloading or evicting a model's weights without starting a veloxquant serve process:

const pulled = await vq.models.pull("mlx-community/Qwen3-4B-4bit");
// { id: "mlx-community/Qwen3-4B-4bit", sizeBytes: 2278969756 }

const deleted = await vq.models.delete("mlx-community/Qwen3-4B-4bit");
// { id: "mlx-community/Qwen3-4B-4bit", freedBytes: 2278969756 }

Both are built on the same huggingface_hub APIs the cache scanner above already relies on: pull() calls snapshot_download(), and delete() uses scan_cache_dir() → delete_revisions() → execute() rather than an rm -rf on a resolved path — the cache's blob layout is content-addressed and shared across revisions/repos via symlinks, so a naive recursive delete risks corrupting a different cached model.

vq.models.pull() does not default to VeloxQuantOptions.timeoutMs's usual 30s CLI timeout — model downloads can take many minutes, so pass { timeoutMs } explicitly if you want a bound; by default the call waits as long as it takes. There's no progress callback in this version: snapshot_download()'s tqdm-based progress doesn't cross the subprocess boundary cleanly, so a pull() call is silent until it resolves or throws.

Both are exposed on the CLI too:

npx veloxquant models pull mlx-community/Qwen3-4B-4bit
npx veloxquant models delete mlx-community/Qwen3-4B-4bit

Hardware-aware memory optimization

const info = await vq.system.info();
// { chip: "Apple M4", unifiedMemoryBytes: ..., veloxquantVersion: "0.71.1", ... }

const estimate = await vq.memory.estimate({ seqLen: 32768, headDim: 128, nLayers: 32 });
// { fp16KvCacheBytes, recommendedMethod, estimatedCompressedBytes, memorySavedBytes, reason }

const picked = await vq.optimize({ profile: "maximum-context", seqLen: 32768 });
// picks a method/bit-width biased toward the long-context band

const recommendation = await vq.recommendModel({ modelClass: "7B", goal: "max_context" });
// chip/ramGb auto-detected from this machine when omitted — pass them explicitly to override

Persistent model sessions

Loading a model with automatic optimization keeps the server process alive across multiple turns instead of spinning one up per call:

const model = await vq.load({ model: "mlx-community/Qwen3-8B-4bit", optimize: "auto" });
const r1 = await model.chat({ prompt: "Hi" });
const r2 = await model.chat({ prompt: "Now summarize that" });
await model.stop();

Cancel an in-flight completion or stream without stopping the loaded model by passing an AbortSignal. The same option is available on agent.run():

const controller = new AbortController();
const stream = model.stream({ prompt: "Write a long story", signal: controller.signal });

for await (const chunk of stream) {
  process.stdout.write(chunk.text);
  if (shouldStop()) controller.abort();
}

// The model process is still available for another request.
await model.stop();

Multi-turn conversations

model.conversation() bookkeeps chat history automatically so you don't have to hand-roll a ChatMessage[] and re-append it after every turn:

const model = await vq.load({ model: "mlx-community/Qwen3-4B-4bit", optimize: "auto" });
const convo = model.conversation({ system: "You are a terse assistant." });

const r1 = await convo.send("What's the capital of France?");
console.log(r1.text); // "Paris."

const r2 = await convo.send("And its population?");
console.log(r2.text); // has full history context

console.log(convo.messages); // read-only view of accumulated ChatMessage[]
convo.reset(); // clears history, keeps the system prompt

If you pass tools to send() and the model calls one, you're responsible for executing it and feeding the result back as a tool-role message yourself — Conversation only handles history bookkeeping, not a second agent loop (see Tool-calling agent for that).

Structured output / JSON mode

Ask the model for JSON back with responseFormat:

const result = await model.chat({
  messages: [{ role: "user", content: "Extract the name and age from: John is 30." }],
  responseFormat: {
    type: "json_schema",
    schema: {
      type: "object",
      properties: { name: { type: "string" }, age: { type: "number" } },
      required: ["name", "age"],
    },
  },
});

console.log(result.json); // { name: "John", age: 30 } — already JSON.parse()'d

{ type: "json_object" } works the same way without a schema, for "just give me valid JSON, any shape."

This is not a hard guarantee. The mlx_lm server this SDK wraps has no native support for OpenAI's response_format field — it's not grammar-constrained decoding under the hood. Under the hood this is a best-effort prompt-injection fallback: formatting instructions (and the schema, for json_schema) are appended to the outgoing messages, and the model's text response is parsed as JSON afterward (markdown code fences are stripped first, since models sometimes wrap JSON in them despite being told not to). result.json is null when responseFormat wasn't requested; when it was requested and the model's output fails to parse, chat() throws a descriptive error instead of silently returning null (which would be indistinguishable from "not requested"). responseFormat is still sent on the wire request too, in case a future server version starts honoring it natively — it's an inert no-op today either way. See issue #14 for the investigation into mlx_lm/server.py behind this.

model.conversation().send() accepts the same responseFormat option.

Benchmarking

vq.benchmark() measures tokens/sec, time-to-first-token, and resident memory (RSS) for a model on this machine — real numbers from your own hardware, not a README claim:

import { VeloxQuant } from "@veloxquant/sdk";

const vq = new VeloxQuant();
const result = await vq.benchmark({ model: "mlx-community/Llama-3.2-1B-Instruct-4bit" });

console.log(result.toMarkdown());
VeloxQuant Benchmark

Model: mlx-community/Llama-3.2-1B-Instruct-4bit
Machine: Apple M4
RAM: 24GB

Tokens/sec: 124.8
TTFT: 237ms

turboquant_rvq resident memory: 1055MB
kivi resident memory: 1052MB
Resident memory reduced: 0%

(Real output from a local run — requires real Apple Silicon hardware and the model already downloaded; there's no way to run this in CI.)

Tokens/sec and TTFT are measured at the SDK boundary from stream() chunk timestamps. The resident-memory numbers are genuine measured RSS of the veloxquant serve subprocess (via ps) — not the accounting-only byte counts vq.memory.estimate() reports (see Known limitations). Because of that, compression is not guaranteed to reduce measured RSS: toMarkdown() reports an increase honestly rather than hiding it, since idle RSS mostly reflects model weights, and KV-cache compression's effect is small or even negative for a short prompt against a small model — measured directly against Llama-3.2-1B-Instruct-4bit above.

OpenAI compatibility

A server started with vq.load() already speaks the OpenAI chat-completions API — this SDK's own chat()/stream() post directly to ${baseUrl}/v1/chat/completions. That means the real openai npm package works against it directly, with no wrapper class:

import OpenAI from "openai";
import { VeloxQuant } from "@veloxquant/sdk";

const vq = new VeloxQuant();
const model = await vq.load({ model: "mlx-community/Llama-3.2-1B-Instruct-4bit", optimize: "auto" });

const client = new OpenAI({ baseURL: `${model.baseUrl}/v1`, apiKey: "local" });
const response = await client.chat.completions.create({
  model: model.model, // the served model id — NOT model.method (that's the compression method name)
  messages: [{ role: "user", content: "Hello" }],
});

Two things to get right, both verified against a real running server:

  • baseURL must include /v1 — the openai client appends /chat/completions itself.
  • model in the request must be the actual served model id — model.model, not model.method (that's the KV-cache compression method name, e.g. "kivi" — passing it as model is rejected with a 404, since the server is pinned to one model per process and treats model as a routing check, not a free label).

A full runnable version (including streaming) is at examples/openai-client.ts. openai is not a dependency of this package — install it yourself if you want to use it this way.

Vercel AI SDK

@veloxquant/sdk/ai-sdk wraps an already-loaded model as a Vercel AI SDK LanguageModelV4, for use with generateText()/streamText():

import { generateText, streamText } from "ai";
import { veloxquant } from "@veloxquant/sdk/ai-sdk";
import { VeloxQuant } from "@veloxquant/sdk";

const vq = new VeloxQuant();
const model = await vq.load({ model: "mlx-community/Llama-3.2-1B-Instruct-4bit", optimize: "auto" });

const result = await generateText({ model: veloxquant(model), prompt: "Hello!" });
console.log(result.text);

await model.stop();

veloxquant() takes an already-running VeloxQuantModel, not a bare model name — createOpenAICompatible() (the AI SDK helper this is built on) is synchronous and needs a baseUrl up front, but getting one requires an async vq.load() first, and an AI SDK provider has no disposal hook where this function could safely call model.stop() for you. Loading (and stopping) the model stays the caller's responsibility, same as everywhere else in this SDK.

ai and @ai-sdk/openai-compatible are optional peer dependencies — install them yourself to use this subpath. A full runnable version (including streaming) is at examples/ai-sdk.ts.

LangChain.js

@veloxquant/sdk/langchain wraps an already-loaded model as a LangChain.js BaseChatModel, for use with invoke(), LCEL chains, and agents:

import { VeloxQuant } from "@veloxquant/sdk";
import { VeloxQuantChatModel } from "@veloxquant/sdk/langchain";

const vq = new VeloxQuant();
const model = await vq.load({ model: "mlx-community/Qwen3-4B-4bit", optimize: "auto" });
const chatModel = new VeloxQuantChatModel(model);

const result = await chatModel.invoke("Explain quantum computing simply.");
console.log(result.content);

await model.stop();

Works in LCEL chains the same way any other BaseChatModel does:

import { ChatPromptTemplate } from "@langchain/core/prompts";
import { StringOutputParser } from "@langchain/core/output_parsers";

const prompt = ChatPromptTemplate.fromMessages([
  ["system", "You are a terse assistant."],
  ["human", "{question}"],
]);
const chain = prompt.pipe(chatModel).pipe(new StringOutputParser());
console.log(await chain.invoke({ question: "What's the capital of France?" }));

Same lifecycle rule as @veloxquant/sdk/ai-sdk: VeloxQuantChatModel takes an already-running VeloxQuantModel, not a bare model name, since loading is async and this SDK has no disposal hook to call model.stop() on your behalf — that stays the caller's responsibility.

Streaming works. .stream() and LCEL chains stream real token deltas via _streamResponseChunks(), backed by the same SSE stream VeloxQuantModel.stream() uses:

const stream = await chatModel.stream("Explain quantum computing simply.");
for await (const chunk of stream) {
  process.stdout.write(chunk.content as string);
}

Tool-call deltas stream too — verified against mlx_lm's server, each streamed tool call arrives as one complete { id, name, arguments } object per SSE event rather than an incrementally-assembled fragment, so tool_call_chunks on each AIMessageChunk are always immediately valid on their own, not partial JSON needing further concatenation.

One asymmetry vs. invoke(): mlx_lm's server does not send a usage object on any streamed SSE event (only on the non-streaming response), so token-usage metadata is unavailable when streaming — llmOutput carries no tokenUsage on the streaming path, unlike _generate().

@langchain/core is an optional peer dependency — install it yourself to use this subpath. A full runnable version (direct invoke() and an LCEL chain, verified against a real running server) is at examples/langchain.ts.

LlamaIndex.TS

@veloxquant/sdk/llamaindex wraps an already-loaded model as a LlamaIndex.TS LLM, usable anywhere LlamaIndex expects one — query engines, agents, Settings.llm:

import { VeloxQuant } from "@veloxquant/sdk";
import { VeloxQuantLLM } from "@veloxquant/sdk/llamaindex";

const vq = new VeloxQuant();
const model = await vq.load({ model: "mlx-community/Qwen3-4B-4bit", optimize: "auto" });
const llm = new VeloxQuantLLM(model);

const response = await llm.chat({ messages: [{ role: "user", content: "Hello!" }] });
console.log(response.message.content);

await model.stop();

Streaming works from day one in this adapter (unlike the LangChain.js adapter's original gap):

const stream = await llm.chat({ messages: [{ role: "user", content: "Hi" }], stream: true });
for await (const chunk of stream) process.stdout.write(chunk.delta);

complete() (LlamaIndex's plain-prompt API) works too, built on the same chat() under the hood via LlamaIndex's own BaseLLM:

const completion = await llm.complete({ prompt: "The capital of France is" });
console.log(completion.text);

Same lifecycle rule as the other two adapters: VeloxQuantLLM takes an already-running VeloxQuantModel, not a bare model name — loading is async and this SDK has no disposal hook to call model.stop() on your behalf.

contextWindow cannot be verified against the served model — veloxquant-mlx has no LLM catalog (see Overview), so there's no source of truth for a model's actual context length here. It defaults to LlamaIndex's own DEFAULT_CONTEXT_WINDOW (3900) unless you pass { contextWindow } to the constructor — override it if the served model's real window matters to your use case (e.g. LlamaIndex sizing its own chunking/retrieval logic off this value).

Text-only, no tool-calling in this adapter. LlamaIndex's tool metadata schema doesn't line up with the OpenAI tools shape this SDK's Agent and wire format use elsewhere, and reconciling the two needs a deliberate design pass — out of scope for this version. A multimodal message's non-text content parts (image/audio) are dropped rather than thrown on, so a mixed text+image message doesn't crash a text-only request outright — but no non-text content is ever sent to the model.

llamaindex is an optional peer dependency — install it yourself to use this subpath. A full runnable version (chat, streaming chat, and complete) is at examples/llamaindex.ts.

Tool-calling agent

vq.agent() loads a model and gives you a single-turn tool-calling loop: register tools, call run(), and the model calls tools, gets their results fed back, and produces a final answer.

const agent = await vq.agent({ model: "mlx-community/Qwen3-4B-4bit", optimize: "auto" });

agent.tool({
  name: "get_weather",
  description: "Get the current weather for a location",
  parameters: {
    type: "object",
    properties: { location: { type: "string", description: "City name" } },
    required: ["location"],
  },
  execute: async ({ location }) => ({ location, tempC: 22, condition: "sunny" }),
});

const result = await agent.run("What's the weather in Tokyo? Do I need an umbrella?");
console.log(result.text);
// "The current weather in Tokyo is 22°C with sunny conditions. ... you likely don't need an umbrella."

await agent.stop();

This is built on real, model-native tool-calling — not prompt-engineered parsing in this SDK. veloxquant serve wraps mlx_lm.server, which checks the model's own tokenizer for has_tool_calling and parses tool_calls out of generation using the tokenizer's own chat template and tool-call tokens. Tool definitions and results use the exact OpenAI tools/tool_calls wire shape (verified against a real running server with mlx-community/Qwen3-4B-4bit) rather than a bespoke schema, so the same tool definitions work whether the request goes through agent.tool(), vq.chat({ tools }) directly, or the raw OpenAI client from OpenAI compatibility.

Not every model supports tool-calling — it depends on the model's tokenizer having a tool-calling chat template. If a model doesn't call your tool as expected, that's the first thing to check.

Scope: single-turn tool calling only (call tools → feed results back → repeat until the model stops calling tools or maxSteps, default 8, is reached). No multi-step planning beyond that loop — a natural follow-up but a separate, larger scope.

A full runnable version is at examples/agent.ts.

MCP tool sources

agent.useMcpServer() connects to an MCP server and registers its tools alongside anything registered with agent.tool() — both share one dispatch loop in run():

const agent = await vq.agent({ model: "mlx-community/Qwen3-4B-4bit", optimize: "auto" });

await agent.useMcpServer({
  name: "filesystem",
  transport: "stdio",
  command: "npx",
  args: ["-y", "@modelcontextprotocol/server-filesystem", "/path/to/allowed/dir"],
});

const result = await agent.run("List the files in the allowed directory.");
console.log(result.text);

await agent.stop(); // also closes any MCP servers useMcpServer() connected itself

Or hand it an already-connected Client from @modelcontextprotocol/sdk directly ({ name, client }) — in that case the connection's lifecycle stays with whoever created it, and agent.stop() will not close it.

Can be called more than once, including after the agent has started running, so a long-lived agent can pick up more tools mid-session. Registering a tool whose name collides with an already-registered one (manual or from another MCP server) throws — same behavior as calling agent.tool() twice with the same name.

@modelcontextprotocol/sdk is an optional peer dependency, loaded lazily only when useMcpServer() is actually called — the rest of this SDK works without it installed. Scope: MCP tools only, no resources or prompts primitives. A tool result with unsupported content (image/audio/resource) throws a clear error rather than silently dropping it, since surfacing non-text content to a text-only chat model needs a deliberate design decision this SDK hasn't made yet.

CLI

npx veloxquant doctor      # checks Apple Silicon + veloxquant-mlx install
npx veloxquant recommend   # shows detected hardware + servable methods
npx veloxquant recommend --model-class 7B --goal everyday  # full recommendation (chip/RAM auto-detected)
npx veloxquant analyze --seq-len 32768 --head-dim 128 --n-layers 32
npx veloxquant serve --model mlx-community/Qwen3-4B-4bit   # persistent local server

vq serve starts a long-running local server from the terminal — useful for pointing a separate frontend, curl, or another language's OpenAI client at it without writing a Node script:

npx veloxquant serve --model mlx-community/Qwen3-4B-4bit --optimize
veloxquant serve ready
  baseUrl: http://127.0.0.1:53021
  model:   mlx-community/Qwen3-4B-4bit
  method:  turboquant_rvq
  pid:     61481

Press Ctrl+C to stop.

(Real output from a local run.) It reuses the same startServer() the SDK's vq.load() calls internally — same VELOXQUANT_READY handshake, same ServeHandle shutdown path — so behavior is identical whether the server is started from JS or from this CLI command. Ctrl+C (or SIGTERM) stops the underlying veloxquant serve subprocess cleanly rather than leaving it orphaned.

--model <id>      Model to serve (required)
--method <name>   KV-cache compression method (default: turboquant_rvq)
--bits <n>        Bit width override
--port <n>        Port to listen on (default: a free port)
--host <addr>     Host to bind (default: 127.0.0.1)
--optimize        Auto-tune method/bits/knobs for detected hardware
--json            Print the ready state as JSON instead of formatted text

Every command accepts --json for scripting/CI use — valid, parseable JSON on stdout with nothing else mixed in:

npx veloxquant doctor --json
{
  "ready": true,
  "platform": { "ok": true, "value": "darwin" },
  "appleSilicon": { "ok": true, "chip": "Apple M4" },
  "python": { "ok": true, "interpreter": "python3" },
  "veloxquantMlx": { "ok": true, "version": "0.70.0" }
}

(Real output, verified against a local machine.)

Known limitations

recommend vs. auto-config stay separate. veloxquant-mlx's recommend CLI (chip/RAM/model-class picker) and auto-config (workload/hardware → method picker) are two independent selectors in the underlying package — this SDK exposes both (vq.recommendModel() and vq.memory / vq.optimize()) rather than merging them, since the package itself doesn't merge them either.

Compression byte counts are accounting-only. Caches store dequantized fp16 tensors today, so byte counters measure compression fidelity, not runtime memory actually freed. memory.estimate() surfaces this as estimatedCompressedBytes / memorySavedBytes — treat them as directional, not a guarantee, until the package changes this (see veloxquant-mlx issue #27).

recommendModel() inputs are constrained to what's installed. chip/ramGb are validated against whatever the installed veloxquant recommend CLI actually accepts (SUPPORTED_CHIPS, SUPPORTED_RAM_GB exported from this package) — M1-M4 and RAM up to 128GB as of veloxquant-mlx 0.71.1. Passing values outside that set fails with the CLI's own argparse error rather than being silently widened.

Development

npm install
npm run build       # tsup -> dist/
npm test            # node --test against test/unit
npm run typecheck
npm run lint

Releasing

Releases and release notes are automated with Changesets. Any PR with a user-facing change (new feature, bug fix, breaking change) should include a changeset:

npx changeset

This prompts for a bump type (patch/minor/major) and a summary, and writes a .changeset/*.md file — commit it with your PR. CI fails PRs that touch src/ without one (use npx changeset add --empty if a change genuinely needs no release notes).

On merge to master, a GitHub Action opens or updates a "Version Packages" PR that bumps package.json and writes CHANGELOG.md from pending changesets. Merging that PR builds, publishes to npm, tags the release, and creates a GitHub Release with the generated notes — no manual changelog writing or tagging required.

License

MIT