@johnhenry/laya-cli
v0.3.1
Published
laya predict / laya bench command line
Maintainers
Readme
@johnhenry/laya-cli
The laya command runs Laya typed decisions from a shell. It works under
Node ≥ 24 and Bun.
Install
npm install @johnhenry/laya-cli
bun add @johnhenry/laya-cliNode ≥ 24 or Bun ≥ 1.2. For GPU inference also install a backend next to it (@johnhenry/backend-mlx on Apple Silicon, @johnhenry/backend-webgpu elsewhere); without one it runs on the CPU reference.
npx @johnhenry/laya-cli predict --state "I was billed twice." --questions @questions.json
bunx @johnhenry/laya-cli bench --backend mlxlaya predict --model aac6fef/laya-mlx \
--state "I was billed twice. Please refund the duplicate." \
--questions '{"department": {"type": "choice", "instructions": "Who should handle this?", "criteria": ["billing", "technical", "sales"]}}'
laya predict --route --state @email.json --questions @questions.json # pick the checkpoint by language
laya bench --backend mlx # latency and throughputlaya predict
laya predict prints the result as JSON on stdout, in the same shape and
formatting as laya-mlx predict (json.dumps(result, ensure_ascii=False,
indent=2); floats print as floats, e.g. 1.0). With --backend mlx the
output is byte-identical to the Python CLI for the README example.
| flag | |
|---|---|
| --model <id\|path> | Hub repo or local checkpoint (default aac6fef/laya-mlx) |
| --state <json\|@file\|text> | parsed as JSON when it is valid JSON, @file reads a JSON file, anything else is plain text |
| --state-file <file> | JSON state file (the laya-mlx flag) |
| --questions <json\|@file\|file> | question definitions |
| --route | choose the checkpoint with @johnhenry/laya-router (the MLX_MODELS repos); the result gains routing. --model then names a checkpoint (english, ml, typed, …); --lang and --task are routing hints |
| --backend auto\|mlx\|webgpu\|cpu | default auto: MLX, then WebGPU, then CPU |
| --dtype f16\|f32 | default f16; float16 and float32 are accepted too |
| --device gpu\|cpu | MLX device |
| --subfolder, --revision, --batch-size (16), --offline | |
Exit codes: 0 on success, 2 for usage errors, 1 for runtime errors (message on stderr). The temperature-clamping warning goes to stderr.
laya bench
laya bench mirrors laya-mlx benchmarks/worker.py. The workload is
laya-mlx's examples/state.json with 1 question (the choice) or 50 questions
(the three example questions, cycled), and --batch-size defaults to 64. It
measures end-to-end predict wall time after --warmup (5) runs, over
--iterations (50) runs, and reports P50 and P95 for one question, throughput
for 50 questions (50 000 / mean ms), load time, peak MLX allocation, and
backend/device/host details. --json prints the report as JSON.
laya quantize
laya quantize --model <repo|dir> --bits 8|4 --out <dir> writes a q8 or q4
copy of a checkpoint as a complete checkpoint directory:
model.safetensors;- configs, tokenizer,
LICENSEandNOTICE, copied from the source; - a
README.mdsaying the copy is derived from the Apache-2.0 Laya checkpoint.
load() from @johnhenry/laya reads the result from a directory, a Hub repo
or any URL, and dequantizes it while loading. q8 halves the download and q4
cuts it to about 28%. Accuracy and sizes for the three published checkpoints
are in docs/RESULTS.md,
and the format is described in docs/QUANTIZATION.md.
laya quantize --model aac6fef/laya-mlx --bits 8 --out ./laya-mlx-q8
laya quantize --model aac6fef/laya-multilingual-mlx --bits 4 --out ./laya-multilingual-q4| flag | |
|---|---|
| --bits 8\|4 | required |
| --out <dir> | required; an existing model.safetensors there is refused without --force |
| --group-size <n\|row> | values per scale along the input dimension (default 64) |
| --no-embeddings | keep the token embedding in float |
| --exclude <regex> | keep matching tensors in float (repeatable) |
| --q8 <regex> | with --bits 4: store matching tensors as q8 (repeatable) |
| --no-refine | q4: plain min/max ranges instead of the least-squares refit |
| --offline, --revision, --subfolder | as for predict |
Quantizing the English checkpoint takes about 10 s (q8) or 30 s (q4), and peaks at about 1.1 GB of memory.
Running from a checkout
bin/laya.js runs dist/bin.js when the package is built. In a workspace
checkout without a build, it re-executes itself under
--conditions=source so the TypeScript sources are used directly.
node packages/laya-cli/bin/laya.js … and bun packages/laya-cli/bin/laya.js …
both work.
Limitations
--statethat happens to be valid JSON (for example123or"quoted") is parsed as JSON. Use--state-fileor@fileto be explicit.laya quantizeoutput is for laya-js only: the Python laya-mlx runtime cannot read quantized files. It reads a Hub repo or a local directory, not a URL.- There is no
convertcommand (laya-mlxconvertwrites MLX checkpoints; use the Python tool).
Family
Part of laya-js, Laya typed decisions in JavaScript on MLX, WebGPU and CPU — see its package map and results.
- A thin shell over
@johnhenry/laya(load/predict) and@johnhenry/laya-router(--route), printing JSON with@johnhenry/pyjsonexactly aslaya-mlx predictdoes.
License
Apache-2.0. Ports logic from laya-mlx and Laya (both Apache-2.0); see NOTICE.
