npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@layerscale/layerscale

v0.7.0

Published

Client for the LayerScale inference server

Readme

@layerscale/layerscale

TypeScript client for the LayerScale inference server. Zero runtime dependencies, built on the global fetch and WebSocket (Node >= 22).

Install

npm install @layerscale/layerscale

Authentication

By default the server is open and needs no key. Started with --api-key, its Ed25519 license key (LSK-...) is the API bearer token: every route except the OPTIONS preflight and the health probes (/health, /healthz) requires Authorization: Bearer <LSK-...> (x-api-key is also accepted) and answers 401 authentication_error without it, and the WebSocket upgrade is refused with a plain-HTTP 401 before the handshake. Pass the key as apiKey or via the LAYERSCALE_API_KEY env var; LAYERSCALE_LICENSE_KEY, the server's own env var for the key, is accepted as a fallback. The client sends it as Authorization: Bearer <key> on every request, the WebSocket handshake included.

Quick start

import { LayerScale } from '@layerscale/layerscale';

// Base URL from LAYERSCALE_BASE_URL, falling back to http://127.0.0.1:8080.
const client = new LayerScale();

const health = await client.health.check(); // never throws
console.log(health.ok, health.status, health.body?.status);

const res = await client.chat.create({
    messages: [{ role: 'user', content: 'Hello!' }],
});
console.log(res.choices[0].message.content);

Client options

const client = new LayerScale('http://127.0.0.1:8080', {
    apiKey: undefined,   // license key (LSK-...); falls back to LAYERSCALE_API_KEY, then LAYERSCALE_LICENSE_KEY
    headers: {},         // extra headers for every request
    timeoutMs: 600_000,  // default timeout for non-streaming requests (10 min); 0 disables
    maxRetries: 2,       // GET-only retries on connect errors / 408 / body-budget 503
});

The base URL falls back to the LAYERSCALE_BASE_URL env var, then to http://127.0.0.1:8080 (the server's default bind).

Every method takes an optional final { signal } with an AbortSignal; it is combined with the default timeout via AbortSignal.any. Streaming calls are exempt from the default timeout, the caller controls their duration through signal.

Endpoints

| Client call | Wire | |---|---| | client.chat.create(params) / client.chat.stream(params) | POST /v1/chat/completions | | client.messages.create(params) / client.messages.stream(params) | POST /v1/messages | | client.models.list() / client.models.retrieve(id) | GET /v1/models / GET /v1/models/{id} | | client.health.check() | GET /health | | client.metrics.fetch() | GET /metrics (Prometheus text) | | client.sessions.create(params?) | POST /v1/sessions | | client.sessions.list() / .get(id) / .delete(id) | GET/GET/DELETE /v1/sessions[/{id}] | | client.sessions.push(id, data) | POST /v1/sessions/{id}/push | | client.sessions.generate(id, params) / .generateStream(...) | POST /v1/sessions/{id}/generate | | client.sessions.flash(id, query, maxTokens?) | POST /v1/sessions/{id}/flash | | client.sessions.listFlash(id) / .unflash(id, fid) | GET / DELETE /v1/sessions/{id}/flash[/{fid}] | | client.sessions.events(id) | GET /v1/sessions/{id}/events (SSE) | | client.sessions.ws(id) | WebSocket /v1/sessions/{id}/ws |

This resource-style surface matches the Python client (client.chat.create(...) ↔ client.chat.create(...), and so on). The pre-0.6 flat methods, client.chat(params), client.chatStream, client.message, client.messageStream, client.models(), client.model(id), client.health(), and client.sessions.stream(id), remain as deprecated aliases.

There is no /v1/completions endpoint and no /v1/health, health lives at /health (or /healthz), outside /v1. client.metrics.fetch() returns the Prometheus text exposition (text/plain; version=0.0.4) verbatim as a string; unlike the health probes, /metrics is not auth-exempt, so under --api-key a scraper needs the license key too.

Sampling defaults (all generation endpoints)

| Parameter | Default | |---|---| | temperature | 0.6 | | top_p | 0.9 | | max_tokens | 512 (required on /v1/messages) |

Unsupported OpenAI/Anthropic parameters (n, penalties, logit_bias, top_k, thinking, non-text response_format, ...) are refused by the server with a 400 naming the parameter, never silently ignored.

Chat completions (OpenAI-compatible)

const res = await client.chat.create({
    messages: [
        { role: 'system', content: 'You are terse.' },
        { role: 'user', content: 'What is a candlestick chart?' },
    ],
    max_tokens: 128,
    temperature: 0.2,
});
console.log(res.choices[0].message.content, res.usage);

model is optional and ignored, the server always answers with its own model name. Tool calling uses the standard tools / tool_choice shape: 'auto', 'none', 'required' or a named function (the last two force a call on Qwen3 and Mistral templates and run the turn without thinking), and parallel_tool_calls: false limits the turn to one call. On a prose reply, message.content is a string and the tool_calls key is absent; with tool calls, content is the text ahead of the first call, or null when there was none. reasoning_effort ('none' … 'xhigh') or chat_template_kwargs.enable_thinking switch a reasoning model's thinking. logprobs: true (plus top_logprobs, 0–20) fills choices[0].logprobs; when streaming, it arrives on the final chunk. User messages may carry image_url parts with base64 data: URLs on a model with image input. usage.prompt_tokens_details.cached_tokens appears only when > 0 (session turns).

Streaming (SSE dialect A)

Chat and session-generate streams are untyped data:-only SSE terminated by a literal data: [DONE], which the client consumes. Heartbeat comment lines are ignored.

for await (const chunk of client.chat.stream({
    messages: [{ role: 'user', content: 'Hi' }],
    stream_options: { include_usage: true },
})) {
    const choice = chunk.choices[0];
    if (choice?.delta.content) process.stdout.write(choice.delta.content);
    if (!choice && chunk.usage) console.log(chunk.usage); // final usage chunk: choices is []
}

Chunks carry model. Each detected tool call arrives as exactly two chunks: one with id/type/function.name, then one with the full function.arguments JSON string. With stream_options: { include_usage: true }, every chunk carries usage: null until a final usage chunk whose choices is [].

Messages (Anthropic-compatible)

const res = await client.messages.create({
    messages: [{ role: 'user', content: 'What is a candlestick chart?' }],
    max_tokens: 128, // REQUIRED on this surface
    system: 'You are terse.', // string or an array of text blocks
});
console.log(res.content, res.stop_reason, res.usage);

Roles are user/assistant only (system is top-level). tool_choice takes auto, none, any or {type:'tool', name} (the last two force a call), and disable_parallel_tool_use: true beside the type limits the turn to one call. top_k and thinking are rejected. User content may carry image blocks with a base64 source on a model with image input. usage.cache_read_input_tokens appears only when > 0.

Streaming (SSE dialect B)

/v1/messages streams are typed event: + data: SSE with no [DONE] sentinel, the stream ends after message_stop:

for await (const event of client.messages.stream({
    messages: [{ role: 'user', content: 'Hi' }],
    max_tokens: 128,
})) {
    if (event.type === 'content_block_delta' && event.delta.type === 'text_delta') {
        process.stdout.write(event.delta.text);
    }
}

Event order: message_start → (content_block_start / content_block_delta* / content_block_stop, text first, then one block per tool call with a single input_json_delta) → message_delta (whose usage includes input_tokens, unlike Anthropic's hosted API) → message_stop.

Sessions

Sessions are explicit, durable KV-cache-backed contexts shared by both chat surfaces. They never expire on their own (max 256 live; a license tier may cap far fewer).

const session = await client.sessions.create({
    id: 'market',      // optional; 1–128 chars of [A-Za-z0-9_-]; server-generated otherwise
    pre_decode: false, // optional
    window: 4096,      // optional sliding window (token count); absent → append-only
});

create answers 403 (permission_error, "the <tier> license permits at most N live session(s); delete one first") when the license tier's session cap is filled, checked before the engine's 256-session backstop, and 409 for a duplicate id or when 256 sessions are live. The session object reports tokens, kv_tokens, busy, turns, data_end, data_version, pending, ready, evicted, and, only when set, frozen_end, window, ready_token, ready_gap.

Bind a chat turn to a session by passing session_id in a chat/message request. Unknown ids are 404 (never auto-created); a second concurrent turn is 409. Resumed KV shows up as cached_tokens / cache_read_input_tokens.

Push (raw text ingestion)

const ack = await client.sessions.push('market', ['line 1', 'line 2']);
// { pushed, dropped, pending, pending_bytes, total_pushed, total_dropped, data_version }

Push takes a string or an array of strings, there are no typed entries. The ack is immediate; ingestion happens in the background and backpressure never blocks or errors (overflow drops the oldest unprocessed entries). Pass { wait: true } to poll GET /v1/sessions/{id} until the batch is actually ingested and KV-resident, that is, until pending has drained to 0 and data_version has advanced past the ack's and the KV covers the durable prefix (kv_tokens + 1 >= data_end, the server's own residency predicate, mirroring the server's test-suite wait):

await client.sessions.push('market', 'context line', { wait: true, waitTimeoutMs: 30_000 });

pending === 0 alone is not proof of ingestion: the server's ingest worker empties the pending queue the moment it claims the batch, before anything is tokenized or prefetched, and a cancelled pass can bump data_version while the KV still trails. One caveat: a batch whose entries are all dropped as inadmissible during the pass never settles a durable turn, so no version bump ever comes, that wait ends only at waitTimeoutMs (the ack's dropped counts push-time overflow drops only).

Generate (ephemeral queries)

Prompt-only queries over the session's durable prefix (the Q/A tail is evicted by the next durable change):

const res = await client.sessions.generate('market', {
    prompt: 'Is the market bullish or bearish?',
    max_tokens: 32,
    fast_answer: ['bullish', 'bearish'], // optional speculative ready-position exit
    gap_threshold: 2.0,
});
console.log(res.text, res.finish_reason, res.data_version, res.usage);

Flash-cache hits add flash: true, flash_id, confidence; speculative hits add speculative: true, logit_gap. Both report zero usage without total_tokens. Streaming (dialect A):

for await (const chunk of client.sessions.generateStream('market', { prompt: '...' })) {
    if (chunk.done) console.log('\n', chunk.usage); // final frame: the response minus text, plus done: true
    else process.stdout.write(chunk.text);
}

The stream ends with data: [DONE] after the done frame; the client consumes it so the connection can be reused.

Flash queries

Standing questions re-evaluated in the background after every ingested batch (max 20 per session):

const q = await client.sessions.flash('market', 'Is the market bullish or bearish?', 16);
// { id, object, query, max_tokens, tokens, fresh }, plus value/data_version/confidence/evaluated_at once answered

const list = await client.sessions.listFlash('market'); // { object, data, data_version }
await client.sessions.unflash('market', q.id);          // { id, object, deleted, remaining }

Duplicate query strings are 409. maxTokens is clamped server-side to 1..=256 (default 32).

Events (SSE)

A typed event stream (dialect B, no [DONE]): a connected event, then one flash_ready replay per cached answer, then live data_updated (on every durable settle) and flash_ready (when an answer changes) events. Heartbeat comments arrive after every 15 s of silence and are ignored by the client.

for await (const event of client.sessions.events('market')) {
    switch (event.type) {
        case 'connected':    console.log(event.session_id, event.data_version, event.flash_queries); break;
        case 'data_updated': console.log(event.data_version, event.tokens, event.pending); break;
        case 'flash_ready':  console.log(event.id, event.query, event.value, event.confidence); break;
    }
}

The stream ends without error on session delete, server shutdown, or when the subscriber falls > 64 events behind, reconnect and resync from the new connected + replay.

WebSocket

const socket = client.sessions.ws('market'); // sessions.stream(id) is a deprecated alias

socket.on('open', () => socket.push('a line of text'));
socket.on('connected', (d) => console.log(d.session_id, d.data_version, d.flash_queries));
socket.on('push_ack', (d) => console.log(d.pushed, d.dropped, d.pending, d.data_version));
socket.on('data_updated', (d) => console.log(d.data_version, d.tokens, d.pending));
socket.on('flash_ready', (d) => console.log(d.query, '→', d.value));
socket.on('error', (e) => console.warn(e.message, e.code)); // JSON-level errors are non-fatal
socket.on('close', () => console.log('closed'));

socket.ping();  // JSON {"type":"ping"} → a 'pong' event
socket.close(); // JSON {"type":"close"}, then the WebSocket closing handshake

The server sends a protocol ping (hb) every 15 s; the global WebSocket implementation answers it automatically. JSON-level error messages leave the connection open. The socket does not auto-reconnect.

A server started with --api-key requires the license bearer on the upgrade and refuses a missing/invalid key with a plain-HTTP 401 before the handshake. undici reports that refusal as a generic error event, surfaced as { message, code: 0 } with no HTTP status, indistinguishable from a network failure (or a 404 for an unknown session). The Authorization header on the upgrade rides on undici's non-standard { headers } WebSocket option (Node >= 22); browsers cannot attach headers to a WebSocket handshake at all, so browser code cannot open the session WebSocket against an --api-key server.

Errors

Every failure the client itself produces is a subclass of LayerScaleError, so one instanceof check catches them all (mirroring except LayerScaleError in the Python client):

| Class | Thrown for | Extra fields | |---|---|---| | LayerScaleError | base class (never thrown directly) |, | | LayerScaleAPIError | any non-2xx HTTP response | status, body, type | | LayerScaleStreamError | an error frame delivered mid-stream (after the 200 status line, so there is no HTTP status) | body, type | | LayerScaleConnectionError | the request never produced a response (connect refused, reset, DNS) | cause | | LayerScaleTimeoutError (extends LayerScaleConnectionError) | the configured timeoutMs elapsed | cause |

A user-initiated abort (your own AbortSignal) is rethrown unwrapped, cancellation is not a client failure.

The Python correspondence: LayerScaleError ↔ layerscale.LayerScaleError, LayerScaleAPIError.status/.type ↔ APIStatusError.status_code/.type, LayerScaleConnectionError/LayerScaleTimeoutError ↔ APIConnectionError/APITimeoutError, LayerScaleStreamError ↔ APIStreamError.

LayerScaleAPIError.body is the server's parsed error envelope:

  • errors generated inside the POST /v1/messages handler, plus the router-level 401 on that route: {"type": "error", "error": {"type", "message"}} (Anthropic shape)
  • everything else, every other route, plus HTTP-layer rejections on any route including /v1/messages (400 framing errors, 408, 413, 431, 501, the body-budget 503) and the router's 405: {"error": {"message", "type"}} (OpenAI shape)

Do not infer the route from the envelope shape, key off status and message instead. Both shapes carry body.error.message (the Error message) and body.error.type (surfaced as .type). The server derives type from the status, OpenAI shape: 401 → authentication_error, 402/403 → permission_error, 400/404/405/408/409/413/431 → invalid_request_error, 429 → rate_limit_error, else server_error; Anthropic shape: 401 → authentication_error, 403 → permission_error, 404 → not_found_error, 429 → rate_limit_error, other 4xx → invalid_request_error, 529 → overloaded_error, else api_error.

import { LayerScaleAPIError, LayerScaleError } from '@layerscale/layerscale';

try {
    await client.sessions.get('nope');
} catch (err) {
    if (err instanceof LayerScaleAPIError) console.error(err.status, err.message, err.body);
    else if (err instanceof LayerScaleError) console.error('transport failure:', err.message);
}

Mid-stream errors are thrown from the stream generators as LayerScaleStreamError. client.health.check() never throws.

Retries: only GET requests are retried, and only on connect errors, 408, and the 503 body-budget response (maxRetries, default 2). POST and DELETE requests are never retried, a retried DELETE whose response was lost would turn an already-completed delete into a spurious 404, and 409 is never retried (on this server it means duplicate-create or session-busy, not a transient conflict). Streams never retry.

Examples

Runnable examples live in example/ (npm run start-http, npm run start-ws); configure via .env (see .env.example): set LAYERSCALE_BASE_URL and the license key (LAYERSCALE_API_KEY, or the server's own LAYERSCALE_LICENSE_KEY); a server run with --api-key answers 401 on everything but the health probes without it.

Development

npm run typecheck            # tsc over src
npm run typecheck:examples   # tsc over src + example
npm run check                # biome check --write
npm test                     # wire-exact tests against a local mock server
npm run build                # dist/