laya-system-one
v1.3.3
Published
Self-contained System 1 decision engine: multilingual INT8 transformer served by a zero-dependency native binary (or the bundled pure-Rust WASM engine). TypeSafe Jev wire-compatible (/v1/systemone). Runs offline in Node.js, Bun and browsers.
Downloads
1,822
Maintainers
Readme
Laya System-One ⚡
A fast, self-contained decision engine. Give it any text and a set of typed questions, and it answers them — offline, in milliseconds, in over 100 languages. Drop-in compatible with the TypeSafe Jev API (
POST /v1/systemone).
Runs entirely on your machine. No Python, no PyTorch, no API keys, no cloud
calls at inference time. One npm install and it works.
Built on Laya by Convai Innovations — a community project. The model is theirs; this package makes it run in Node.js, Bun and the browser with no Python in the loop. See Credits.
🌟 Why Laya System-One?
- 🔒 100% Offline: Nothing leaves your machine. Ideal for corporate intranets, edge servers and privacy-sensitive workflows.
- ⚡ Fast: ~40 ms per question on a warm engine, measured on every platform we ship for.
- 🔄 TypeSafe Jev Compatible: Drop-in
POST /v1/systemone. Point an existing Jev client at it and it just works. - 🌍 Multilingual: Understands English, Portuguese, Spanish, German, French, Chinese, Japanese and 100+ more, out of the box.
- 💻 Node.js, Bun and Browsers: Native binary on Node and Bun, WebAssembly in the browser.
- 📦 Zero Dependencies:
dependenciesis empty. Nothing to compile, nothing to install system-wide, nothing to keep patched. - 🧩 Two Ways to Run It: As a local HTTP service via the CLI, or in-process for zero network overhead.
📦 Installation
# npm
npm install laya-system-one
# bun
bun add laya-system-one
# pnpm
pnpm add laya-system-oneThe right engine for your machine is installed automatically. The model (~324 MB) is fetched once on first use and cached.
🚀 Quick Start
1. Launch the HTTP service
npx laya-system-one --port 8080| Flag | Env | Default | Description |
| :--- | :--- | :--- | :--- |
| --port <number> | PORT | 8080 | Port to bind |
| --host <string> | HOST | 0.0.0.0 | Address to bind |
| --backend <type> | LAYA_BACKEND | native | native or wasm |
| --api-key <token> | LAYA_API_KEY | (none) | Require Bearer auth on /v1/systemone |
With authentication:
npx laya-system-one --port 8080 --api-key secret-token-xyz2. Use it in-process (zero network overhead)
import { Laya } from 'laya-system-one';
// 1. Initialize the engine
const laya = await Laya.load();
// 2. Define the state (string, object, or array)
const state = {
customer_id: 'cust_9821',
message: 'We were charged twice on our March invoice. Please refund the duplicate amount or we will cancel our plan.'
};
// 3. Define typed questions
const questions = {
department: {
type: 'choice',
instructions: 'Which team should resolve this customer inquiry?',
criteria: {
billing: 'Invoices, refunds, and duplicate charges',
tech_support: 'Software bugs, outages, and error messages',
sales: 'Upgrades, plan changes, and enterprise contracts'
}
},
urgency: {
type: 'score',
instructions: 'Assess the urgency level of this inquiry.',
criteria: ['Low / routine', 'Moderate', 'Critical / blocking / angry']
},
churn_risk: {
type: 'noul',
instructions: 'Does this message present an explicit risk of customer churn?',
threshold: 0.5
}
};
// 4. Evaluate
const result = await laya.predict(state, questions);
console.log(result.answers.department.choice); // -> "billing"
console.log(result.answers.department.confidence); // -> 1.0
console.log(result.answers.urgency.score); // -> 1.95
console.log(result.answers.churn_risk.noul); // -> 0.968
console.log(result.answers.churn_risk.decision); // -> true3. Serve it from inside your app
import { serve } from 'laya-system-one';
const srv = await serve({ host: '127.0.0.1', port: 8080, apiKey: 'optional-key' });
console.log(`Laya server running at ${srv.url}/v1/systemone`);
// later:
await srv.close();In the browser
The wasm backend runs in a browser. Laya.load() picks it automatically
there — native spawns a process, which a browser cannot do — and the model,
tokenizer and wasm engine are fetched over HTTP:
<script type="module">
import { Laya } from 'https://esm.sh/laya-system-one';
const laya = await Laya.load({
modelDir: '/models/', // where model.onnx lives
wasmBase: 'https://cdn.example/laya/wasm/' // optional: wasm from a CDN
});
const out = await laya.predict('We were billed twice and want a refund.', {
department: {
type: 'choice',
instructions: 'Which department should handle this?',
criteria: { billing: 'refunds', tech: 'bugs', sales: 'upgrades' }
}
});
console.log(out.answers.department.choice); // "billing"
</script>Serve models/model.onnx (324 MB) and the src/wasm-pkg/ directory over HTTP
with the right content types (.wasm as application/wasm), and enable
cross-origin isolation if you want the threaded build. The model is fetched
once and can be cached by the browser like any other asset.
Two honest notes:
- The wasm backend is roughly 40x slower than the native binary — a question takes seconds in a browser, not milliseconds. It exists so the browser works at all.
- The environment detection, the HTTP fetching of the model/tokenizer/wasm
bytes and the inline engine are covered by tests on Node. The one step that
cannot be — importing the wasm glue over
http:— is refused by Node's ESM loader, so it is exercised in a browser rather than in CI.
Laya.load(options) options
| Option | Default | Description |
| :--- | :--- | :--- |
| backend | 'native' | native (bundled Rust server) or wasm |
| modelDir | the package's models/ | where model.onnx and tokenizer.json live |
| maxLen | 2048 | token budget for the state — raise it for long documents (max 8192, see Long inputs) |
| apiKey | null | require a Bearer token on the HTTP layer |
| port / host | 0 / 127.0.0.1 | where the native server binds |
| threads | 0 | inference threads (0 = runtime default) |
📡 HTTP API Reference (TypeSafe Jev compatible)
POST /v1/systemone
Host: localhost:8080
Content-Type: application/json
Authorization: Bearer <API_KEY> [optional unless configured]| Parameter | Type | Required | Description |
| :--- | :--- | :--- | :--- |
| state | string | object | array | Yes | The context or text being evaluated. |
| questions | Record<string, Question> | Yes | Map of question keys to typed questions. |
| model | string | No | Model name (defaults to laya-multilingual, echoed back). |
Question types
choice — pick one of several options:
{
"type": "choice",
"instructions": "Which department should handle this ticket?",
"criteria": {
"billing": "Invoices and credit card transactions",
"technical": "Software bugs and service disruptions"
}
}score — place on an ordered scale:
{
"type": "score",
"instructions": "Rate the severity of the issue.",
"criteria": ["Minor cosmetic issue", "Degraded functionality", "Critical full service outage"]
}noul — calibrated yes/no probability:
{
"type": "noul",
"instructions": "Does the user explicitly request a refund?",
"threshold": 0.6
}Example request
curl -X POST http://localhost:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"state": { "text": "Fui cobrado duas vezes na minha fatura. Reembolsem imediatamente." },
"questions": {
"dept": {
"type": "choice",
"instructions": "Which department should respond?",
"criteria": { "billing": "Refunds, invoices, and payments", "support": "Technical and product questions" }
},
"urgency": {
"type": "score",
"instructions": "Urgency rating",
"criteria": ["Low", "Medium", "High"]
},
"refund_demanded": {
"type": "noul",
"instructions": "Is the customer requesting a refund?",
"threshold": 0.5
}
}
}'Example response
{
"model": "laya-multilingual",
"answers": {
"dept": {
"type": "choice",
"choice": "billing",
"probabilities": { "billing": 1.0, "support": 0.0 },
"confidence": 1.0
},
"urgency": {
"type": "score",
"score": 1.9482,
"legend": { "0": "Low", "1": "Medium", "2": "High" },
"probabilities": { "0": 0.0011, "1": 0.0496, "2": 0.9493 },
"confidence": 0.9493
},
"refund_demanded": {
"type": "noul",
"noul": 0.9852,
"confidence": 0.9852,
"threshold": 0.5,
"decision": true
}
},
"usage": { "input_tokens": 82, "output_tokens": 12 }
}Healthcheck
GET /health{
"status": "ok",
"model": "laya-multilingual",
"version": "1.1.0",
"protocol": "TypeSafe Jev /v1/systemone compatible"
}Error codes
401 Unauthorized— API key configured, header missing or wrong.422 Unprocessable Entity— invalid JSON, orstate/questionsmissing.404 Not Found— unknown route.413 Payload Too Large— body over 4 MB.
⚙️ Backends
| Backend | How it runs | When to use |
| :--- | :--- | :--- |
| native (default) | A self-contained Rust server bundled with the package | The normal choice. Fastest, nothing to install. |
| wasm | Pure Rust compiled to WebAssembly, also bundled | Browsers, or platforms with no native build. |
Both ship inside the package — nothing is compiled or downloaded at install
time. Switch with --backend wasm or LAYA_BACKEND=wasm.
📊 Performance
Measured on real hardware, on every platform we ship a binary for, with the
default native backend. Run npm run bench to measure your own machine.
| Platform | Load | First answer | Warm (4 q/call) | Per question | | :--- | ---: | ---: | ---: | ---: | | macOS arm64 | 1.3 s | 268 ms | 220 ms | 51 ms | | Windows arm64 | 1.4 s | 231 ms | 178 ms | 47 ms | | Linux arm64 | 1.6 s | 184 ms | 145 ms | 40 ms | | Windows x64 | 1.7 s | 175 ms | 149 ms | 40 ms | | Linux arm64 (musl) | 1.9 s | 247 ms | 154 ms | 36 ms | | Linux x64 (musl) | 2.3 s | 387 ms | 253 ms | 59 ms | | Linux x64 | 2.5 s | 331 ms | 264 ms | 63 ms | | macOS x64 | 2.9 s | 387 ms | 360 ms | 84 ms |
Load is reading the model into memory. First answer includes warmup. The
warm numbers are sustained latency. Shared CI runners vary by ~20% between
runs, so treat these as orders of magnitude rather than exact figures.
The wasm backend is roughly 40x slower — it exists so browsers and unusual
platforms work at all, not for throughput. It pads prompts to the smallest of
a fixed set of sequence lengths rather than one large size, which is worth
about 4x on short inputs (a 43-token prompt was being padded to 256).
Long inputs
The model reads up to 8,192 tokens, but it ships with a conservative 2,048-token budget so it stays usable on weak machines. The budget is a cap, not a cost: short inputs are unaffected by raising it — a 74-token question answers in ~90 ms whatever the limit is, because the work follows the input's real length.
Raise it when your inputs are long documents:
LAYA_MAX_LEN=8192 npx laya-system-one --port 8080const laya = await Laya.load({ maxLen: 8192 });Measured on one machine (Windows arm64, native backend), by input length:
| tokens | default (2048) | maxLen: 8192 |
| ---: | ---: | ---: |
| 74 | 88 ms | 88 ms |
| 1,000 | 0.9 s | 0.9 s |
| 2,000 | 6.4 s | 6.4 s |
| 4,000 | 9.2 s (truncated) | 21.3 s |
| 8,000 | 9.2 s (truncated) | 190 s |
Two things worth knowing before you raise it:
- Accuracy degrades with length. Upstream measured 16–18 of 20 requests correct up to about 4,000 tokens, and 8–17 of 20 beyond that. Check your own data — long-document accuracy is not something to assume.
- Cost grows steeply. Past ~2,000 tokens the time climbs faster than the input does (attention is quadratic). 8,000 tokens is minutes, not seconds, on a CPU. If you routinely handle documents that long, truncate them yourself to the part that matters, or run the upstream Python package on a GPU.
Truncation is the real risk of leaving it at the default: a long message gets
cut off and the answer can be wrong rather than slow. On a ~3,000-token input
the shipped default answered sales where the full text answers billing.
🧾 Environment variables
| Variable | Effect |
| :--- | :--- |
| LAYA_BACKEND | native or wasm |
| LAYA_MAX_LEN | token budget for the state (default 2048, max 8192) |
| LAYA_MODEL_PATH | Use a model.onnx you already have (file or directory) |
| LAYA_MODEL_CHUNKS_DIR | Directory holding the model chunks |
| LAYA_MODEL_URL | Override where the model is downloaded from |
| LAYA_CACHE_DIR | Where the model is cached |
| LAYA_PREFETCH_MODEL | 1 = download the model during npm install |
| LAYA_SKIP_MODEL_DOWNLOAD | 1 = never download, never prompt |
| LAYA_API_KEY / API_KEY | Require Authorization: Bearer <key> |
| LAYA_SERVE_BIN | Use a specific laya-serve binary |
Offline or air-gapped:
LAYA_PREFETCH_MODEL=1 npm install laya-system-one # fetch during install
LAYA_MODEL_PATH=/opt/models/model.onnx # or bring your own copy
LAYA_MODEL_CHUNKS_DIR=/opt/models/chunks # or a directory of chunks💻 Requirements
| | |
| :--- | :--- |
| Node.js | ≥ 18.17 |
| Bun | ≥ 1.0 |
| Browsers | The wasm backend |
| OS | Linux (glibc and musl/Alpine), macOS (arm64 and x64), Windows (x64 and arm64) |
| Docker | Debian, Ubuntu, Alpine |
No runtime dependencies. The right native binary for your machine is installed automatically — nothing to compile, no system packages to add.
📥 What gets installed
The package itself is small; the heavy parts arrive as dependencies npm picks for your platform, so you only download what you can run.
| | Size |
| :--- | ---: |
| laya-system-one (code, tokenizer, wasm engine) | ~8.5 MB |
| The one native binary for your platform | 8–26 MB |
| The 13 model chunks | ~235 MB total |
| The model on disk, after the first run | ~324 MB |
The model is written next to the package when that directory is writable, and
to your user cache otherwise — so npm i -g and read-only containers work
without extra configuration. Every copy is checksum-verified, and a run that
is killed mid-download leaves nothing corrupt behind.
🛠️ Development
npm install
npm run check # lint, unit, packaging, integration, e2e + install rehearsal
npm run check:quick # the same minus the model-backed suites
npm run test:rehearsal # install from a local registry and use it, Node + Bun
npm run bench # measure on this machine
npm run lint # syntax + packaging + docs consistencyCI builds every platform, proves each binary answers 10 questions, and uploads
the packages as artifacts. verify-published.yml installs a published version
from the real registry on every platform — Node and Bun, including Alpine for
musl — and runs a real inference.
🙏 Credits
This package would not exist without Laya. It is a community project by Convai Innovations — the model, the architecture, the training method and the wire protocol are all theirs. What this package adds is a way to run it where Python is not an option: Node.js, Bun and the browser.
| | |
| :--- | :--- |
| Upstream project | NandhaKishorM/laya |
| Model | convaiinnovations/laya-multilingual (mmBERT-base, 322M params) |
| Other checkpoints | convaiinnovations/laya (English), laya-typed-decisions |
| Demo | Hugging Face Space |
| Method | RLCD — reinforcement learning against strictly proper scoring rules |
| License | Apache-2.0 (upstream and this package alike) |
If you find this useful, the credit belongs upstream — star their repository and consider supporting the author.
📄 License
Apache-2.0 — the same license as the upstream project.
- This package: Italo Almeida — laya-system-one
- Model & upstream: Convai Innovations — NandhaKishorM/laya
