npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@r4ai/laya-web

v0.2.0

Published

Laya typed decisions in browsers and Node.js with ONNX Runtime Web

Downloads

650

Readme

@r4ai/laya-web

License GitHub Pages npm version

Client-side inference runtime for Laya-MLX decision models in browsers and Node.js, powered by ONNX Runtime Web (WebGPU and WebAssembly SIMD).

Laya evaluates structured decisions (typed decisions) directly over input state or text without generating free-form text:

  • Categorical Choice (choice): Selects the best option from discrete candidate labels
  • Ordinal Scoring (score): Evaluates input against an ordered rubric or severity scale
  • Proposition Verification (noul): Computes the probability that a statement holds true

Live Demo

Key Features

  • 100% Client-Side Inference: Evaluates decisions entirely in the browser without sending input data to external servers
  • Web Worker Compatible: Runs smoothly in dedicated Web Workers to keep the UI thread responsive at 60fps
  • WebGPU with Wasm Fallback: WebGPU hardware acceleration with automatic fallback to single-threaded SIMD WebAssembly
  • CPU Embedding Slicing: Slices FP16 embeddings on the CPU to bypass browser WebGPU storage buffer limits on large vocabularies (>128k tokens)
  • Calibrated Decision Confidence: Computes normalized Shannon entropy ($0.0$ to $1.0$) for reliable uncertainty filtering

Architecture

The library runs in both main threads and Web Workers. Hosting the agent inside a dedicated Web Worker isolates heavy tensor computation from the UI:

flowchart TD
    subgraph Host ["Browser Environment"]
        subgraph Main ["Main Thread (UI)"]
            App["Web Application"]
        end

        subgraph Worker ["Web Worker (Recommended)"]
            Agent["Agent (@r4ai/laya-web)"]
            Tokenizer["Tokenizer (JavaScript)"]
            Slicer["CPU Embedding Slicer<br/>(FP16 to FP32)"]
            ORT["ONNX Runtime Web"]
        end
    end

    subgraph Hardware ["Hardware Acceleration"]
        WebGPU["WebGPU (Primary)"]
        Wasm["Wasm SIMD (Fallback)"]
    end

    App -->|"postMessage(state, questions)"| Agent
    Agent --> Tokenizer
    Tokenizer -->|"Token IDs"| Slicer
    Slicer -->|"Token Embeddings"| ORT
    ORT --> WebGPU
    ORT -.->|"Fallback"| Wasm
    ORT -->|"Logits"| Agent
    Agent -->|"postMessage(Prediction)"| App

Installation

npm install @r4ai/laya-web onnxruntime-web
# or
pnpm add @r4ai/laya-web onnxruntime-web
# or
yarn add @r4ai/laya-web onnxruntime-web

Quickstart

Run inference inside a Web Worker to keep the UI thread responsive. Serve model assets and ONNX Runtime Web Wasm binaries from your static file server (see the Vite example).

import { load } from "@r4ai/laya-web";

// 1. Initialize runtime and load model assets
const agent = await load({
  modelUrl: "/models/laya/",
  backend: "auto", // "webgpu" | "wasm" | "auto"
  wasmPaths: "/ort/",
  onProgress: (progress) => console.log("Load progress:", progress),
});

try {
  // 2. Evaluate typed decisions over input context
  const result = await agent.predict(
    "I was charged twice for my subscription this month. Please issue a refund.",
    {
      department: {
        type: "choice",
        instructions: "Which support department should handle this request?",
        criteria: ["Billing & Refunds", "Technical Support", "Sales"],
      },
      urgency: {
        type: "score",
        instructions: "How urgent is this ticket?",
        criteria: ["Normal", "Elevated", "Immediate"],
      },
      refund_requested: {
        type: "noul",
        instructions: "Does the user explicitly demand a refund?",
      },
    },
  );

  console.log("Active backend:", agent.backend);

  // Inspect categorical choice
  const department = result.answers.department;
  if (department.type === "choice") {
    console.log("Department:", department.choice);
    console.log("Probabilities:", department.probabilities);
    console.log("Confidence:", department.confidence);
  }

  // Inspect ordinal score
  const urgency = result.answers.urgency;
  if (urgency.type === "score") {
    console.log("Urgency score:", urgency.score);
    console.log("Score legend:", urgency.legend);
  }

  // Inspect binary verification
  const refund = result.answers.refund_requested;
  if (refund.type === "noul") {
    console.log("P(True):", refund.noul);
    console.log("Confidence:", refund.confidence);
  }
} finally {
  // 3. Release ONNX sessions and memory buffers
  await agent.dispose();
}

Node.js

Node.js 22 or later can use the same ESM import. Conditional exports select the Node runtime, where auto and wasm use single-threaded WASM on the CPU. webgpu is browser-only. No extra native runtime dependency or wasmPaths setup is required.

import { load } from "@r4ai/laya-web";

const nodeAgent = await load({ modelUrl: "./models/laya" });
try {
  console.log(
    await nodeAgent.predict("Please refund the duplicate charge.", {
      refund: {
        type: "noul",
        instructions: "Does the customer request a refund?",
      },
    }),
  );
} finally {
  await nodeAgent.dispose();
}

modelUrl accepts a relative path (resolved against the working directory), an absolute path, a file: URL string, or an HTTP(S) URL. Use the same exported model files listed below. File reads retain progress reporting, manifest size checks, and AbortSignal cancellation. Inference runs locally; keep the model loaded across requests and dispose it when finished. Loading the real checkpoint still requires memory for the assets and the WASM session.

Decision Types

| Type | Target Task | criteria Schema | Output Structure | | :----------- | :---------------------------- | :------------------------------------------------------------------- | :----------------------------------------------------------------------------------------------- | | choice | Categorical classification | String array ["A", "B"] or dictionary { [label]: "description" } | Selected label (choice), per-label distribution (probabilities), and confidence | | score | Ordinal grading on a rubric | Ordered array of level criteria ["low", "medium", "high"] | Expected score (score as a float), level mapping (legend), probabilities, and confidence | | noul | Binary statement verification | Optional { false?: "description", true?: "description" } | Probability of truth (noul as $P(\text{true})$) and decision confidence |

Confidence Calibration

Every answer returns a normalized confidence value in $[0.0, 1.0]$ representing probability concentration rather than ground-truth correctness. Computation depends on question type:

Categorical Choice & Ordinal Score (choice, score)

  • Single Option ($K < 2$): Returns fixed 1.0
  • Multiple Options ($K \ge 2$): Normalized Shannon entropy over option count $K$

$$H(P) = -\sum_{i=1}^{K} P(i) \ln P(i)$$

$$\text{confidence} = \max\left(0, \min\left(1, 1 - \frac{H(P)}{\ln K}\right)\right)$$

  • Yields $1.0$ when probability concentrates entirely on a single option
  • Scales down to $0.0$ when probability distributes uniformly across all options

Binary Verification (noul)

Evaluated as the binary classification margin over calibrated probabilities:

$$\text{confidence} = \max(P(\text{true}), 1 - P(\text{true}))$$

  • Range spans $[0.5, 1.0]$
  • $0.5$ represents maximum ambiguity (equal probability)
  • $1.0$ represents probability concentrated entirely on one outcome

[!NOTE] confidence measures probability distribution sharpness. It does not guarantee prediction correctness or ground-truth accuracy.

Model Assets & Hosting

Deploy the following files under your static modelUrl directory:

| File | Approximate Size | Purpose | | :------------------- | :--------------- | :--------------------------------------------------------------------- | | config.json | ~1.1 KB | Architecture parameters, calibration temperatures, and SHA-256 digests | | model.onnx | ~5.37 MB | ONNX computation graph structure without weights | | model.onnx.data | ~501.20 MB | External model weights downloaded upfront for ONNX Runtime Web | | embeddings.f16.bin | ~393.22 MB | Raw FP16 token embedding table read by CPU memory | | tokenizer/ | ~34.36 MB | Hugging Face tokenizer configuration and vocabulary files |

Export these files locally from the pinned checkpoint:

pnpm install --frozen-lockfile
pnpm model:export

Local Development & Demo

To launch the SolidJS demo application locally:

pnpm install --frozen-lockfile
pnpm model:export
pnpm dev

Open http://127.0.0.1:5173/ in your browser.

Documentation

License & Attribution

Distributed under the Apache-2.0 License.

See NOTICE for copyright notices and third-party attributions.