npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@k2b/nessi

v0.16.0

Published

Minimal agent loop and provider adapters for nessi.

Readme

@k2b/nessi

Minimal agent loop and provider adapters for TypeScript.

Use the package root for the managed nessi() loop with tools, storage, and loop metadata. Use @k2b/nessi/ai when an app only needs provider calls through one normalized message and stream API.

@k2b/nessi replaces the deprecated @valentinkolb/nessi package. Existing APIs and subpaths are unchanged; replace the package scope in dependencies and imports.

Quick start

bun add @k2b/nessi
import { nessi, defineTool, memoryStore } from "@k2b/nessi";
import { ollama } from "@k2b/nessi/ai";
import { z } from "zod";

const weather = defineTool({
  name: "weather",
  description: "Return a fake weather response",
  inputSchema: z.object({ city: z.string() }),
}).server(async ({ city }) => {
  return { city, forecast: "sunny" };
});

const loop = nessi({
  loopId: crypto.randomUUID(),
  provider: ollama("llama3.1", {
    baseURL: "http://localhost:11434",
  }),
  systemPrompt: "You are concise.",
  input: "How is the weather in Berlin?",
  store: memoryStore(),
  tools: [weather],
  temperature: 0,
  maxOutputTokens: 512,
});

const textBlocks = new Set<string>();

for await (const event of loop) {
  if (event.type === "block_start" && event.kind === "text") {
    textBlocks.add(event.blockId);
  }
  if (event.type === "block_delta" && textBlocks.has(event.blockId)) {
    process.stdout.write(event.delta);
  }
  if (event.type === "issue") {
    console.error(event.issue.kind, event.issue.message);
  }
  if (event.type === "loop_end") {
    console.log(event.loopId);
    console.log(event.aggregate.usage);
    console.log(event.aggregate.timing);
  }
}

Every outbound event from one nessi() run carries the same loopId. Pass your own loopId to align events with a persisted request or UI response group, or let Nessi generate one when omitted.

turn_end reports each internal provider turn and closes every turn_start, also when a turn fails or is aborted. A turn that fails or is aborted while the model is still answering carries the partial message with stopReason "error" or "interrupted"; an abort during tool execution carries the stored message. Usage the provider already reported counts in the aggregate and against the credit store. Partial content is stored, except for a context overflow, which is retried after compaction or ends the loop. When the provider itself stops an answer (finishReason: "error", e.g. a content or safety filter), its tool calls are not executed, not even by a later resume, and the loop ends with "error". The final loop_end event includes aggregate, which groups assistant turns, executable tool calls, tool results, validation/execution errors, malformed or cancelled tool streams, summed usage, and timing for the complete logical loop. aggregate.timing.totalElapsedMs is model generation plus active tool execution; approval/client-tool waits are tracked separately as aggregate.timing.actionWaitMs. Helper exports such as mergeUsage(), cloneLoopAggregate(), and mergeLoopAggregates() are available from @k2b/nessi.

Dynamic tools

Pass a resolver when the active tools depend on application state:

const loop = nessi({
  provider,
  systemPrompt,
  input,
  store,
  tools: async () => toolRegistry.activeFor(userId),
});

Nessi resolves dynamic tools before every provider turn. One copied and validated snapshot supplies both the provider schemas and all tool execution for that turn. Changes made during tool execution therefore become visible on the following provider turn, while calls already emitted by the provider remain executable from their original snapshot. Duplicate names and resolver failures end the loop with a runtime_error issue.

Pending tool calls restored from history resolve one fresh snapshot before execution because an in-memory snapshot cannot survive a restart. Keep the resolver side-effect free; persistence, discovery, authorization, and cleanup remain application-owned.

Server tools receive the provider tool call ID through ctx.callId:

const inspect = defineTool({
  name: "inspect",
  description: "Inspect an item",
  inputSchema: z.object({ id: z.string() }),
}).server(async ({ id }, ctx) => {
  return inspectItem(id, { callId: ctx.callId, signal: ctx.signal });
});

nessi.structured() accepts the same pattern for server tools:

const result = await nessi.structured({
  provider,
  input: "Resolve the current task.",
  output: taskSchema,
  tools: () => toolRegistry.activeServerTools(),
});

A structured tool resolver always selects tool_loop mode, even when a snapshot is empty, because later turns may add tools. Nessi appends its internal submit_result tool to every snapshot. Client tools, approval tools, and a user-defined submit_result remain unsupported. A static empty array keeps the direct native/fallback structured-output path.

Historical tool results

Verbose tool output can remain fully persisted without being sent to the model in every later loop. A tool may derive a compact historical representation once after its output passes validation:

const shell = defineTool({
  name: "shell",
  description: "Run a shell command.",
  inputSchema: z.object({ command: z.string() }),
  outputSchema: z.object({
    exitCode: z.number(),
    stdout: z.string(),
    changedFiles: z.array(z.string()),
  }),
  toHistoricalResult: ({ output }) => ({
    exitCode: output.exitCode,
    changedFiles: output.changedFiles,
    excerpt: output.stdout.slice(0, 500),
  }),
}).server(runShell);

Nessi persists the full result and the optional historicalResult together. Provider calls in the originating loop, including resumes with the same loopId, receive the full result. Calls from a different loop receive the historical value instead. Stored messages, events, and loop aggregates remain full and inspectable. Returning undefined skips the historical representation.

If derivation throws, Nessi stores the full successful result, emits a non-fatal tool_historical_result_error issue, and continues the loop. Legacy messages without historicalResult remain unchanged. maxToolResultChars, when set, runs after this selection as the final context-size boundary.

Steering

Use loop.steer() when the process handling new input owns the running loop:

loop.steer("Skip deployment and only prepare the migration.");

For loops running in another worker or process, provide a steering callback that reads pending messages from application-owned persistence:

const loop = nessi({
  provider,
  systemPrompt,
  store,
  input,
  tools,
  steering: ({ loopId, signal }) => steeringQueue.takePending(loopId, { signal }),
});

The callback may return one message, an ordered array, or undefined. Nessi checks it before provider calls and before a normal loop completion. Applied messages are persisted as user messages and emit the same steer_applied event as loop.steer(). The application owns persistence, claiming, and delivery semantics; Nessi only controls when steering can affect the loop.

Structured output

Use nessi.structured() when an app wants a schema-valid typed result instead of a streamed chat response:

import { nessi } from "@k2b/nessi";
import { openrouter } from "@k2b/nessi/ai";
import { z } from "zod";

const result = await nessi.structured({
  provider: openrouter("openai/gpt-4.1-mini", {
    apiKey: process.env.OPENROUTER_API_KEY,
  }),
  input: "Extract a task card for: Ship the onboarding flow by Friday.",
  outputName: "task_card",
  output: z.object({
    title: z.string(),
    due: z.string().nullable(),
    priority: z.enum(["low", "medium", "high"]),
  }),
  temperature: 0,
});

console.log(result.output.title);
console.log(result.structuredMeta);
console.log(result.aggregate.usage);

For providers and schemas that are safe for native structured output, Nessi passes a provider-specific responseFormat. Otherwise it falls back to schema instructions and one repair attempt. input can be a string, content parts, or a full user message, including image file parts when the provider supports images.

nessi.structured() can use server tools for bounded task work. It adds an internal submit_result tool and returns after that tool receives a valid schema value. Client tools, approval tools, and interactive tool bridges remain the job of the full nessi() loop.

Provider-only usage

import { openrouter } from "@k2b/nessi/ai";

const provider = openrouter("openai/gpt-4.1-mini", {
  apiKey: process.env.OPENROUTER_API_KEY,
});

const result = await provider.complete({
  systemPrompt: "Be concise.",
  messages: [
    {
      role: "user",
      content: [{ type: "text", text: "Summarize this package." }],
    },
  ],
});

console.log(result.message.content);

Provider streams use the same block events as the root loop:

const textBlocks = new Set<string>();

for await (const event of provider.stream({ messages })) {
  if (event.type === "block_start" && event.kind === "text") {
    textBlocks.add(event.blockId);
  }
  if (event.type === "block_delta" && textBlocks.has(event.blockId)) {
    process.stdout.write(event.delta);
  }
  if (event.type === "block_end" && event.block.type === "tool_call") {
    console.log("tool call", event.block.name, event.block.args);
  }
  if (event.type === "issue") {
    console.error(event.issue.kind, event.issue.message);
  }
}

Reasoning effort

Set reasoningEffort to control how much a reasoning model thinks. Set it on the provider as a default, or per call on complete(), stream(), nessi() or nessi.structured(). A call value wins over the provider default, and nothing is sent when neither is set, so the model keeps its own default.

const provider = openai("gpt-6.1-sol", { reasoningEffort: "medium" });

const loop = nessi({ provider, systemPrompt, input, store, reasoningEffort: "high" });
const quick = await provider.complete({ messages, reasoningEffort: "none" });

The value is passed through unchanged, so new levels work without a Nessi update. "none" turns reasoning off. Common levels are "none", "minimal", "low", "medium", "high", "xhigh" and "max". Which levels a model accepts is up to the model; unsupported values come back as provider errors.

| Provider | Sent as | |---|---| | openai, vllm, openAICompatible() | reasoning_effort | | openrouter | reasoning: { effort } | | ollama | think; "none" sends false | | anthropic | thinking: { type: "adaptive" } plus output_config.effort; "none" sends thinking: { type: "disabled" } | | gemini | thinkingConfig.thinkingLevel; "none" sends thinkingBudget: 0 | | mistral | reasoning_effort |

Recent vLLM versions also pass the level to the chat template and turn off thinking for templates such as Qwen's when it is "none". On older versions, use extraBody: { chat_template_kwargs: { enable_thinking: false } }. Anthropic models that cannot turn thinking off reject "none"; models before Claude 4.6 need the budget mode through extraBody: { thinking: { type: "enabled", budget_tokens: 4096 } } instead of reasoningEffort. Thinking counts against max_tokens, so the Anthropic default maxOutputTokens is 8192. Gemini 3 cannot turn thinking off completely; use "minimal" there. Gemini 2.5 takes token budgets instead of levels, which extraBody: { generationConfig: { thinkingConfig: { thinkingBudget: 2048 } } } sets. Gemini returns thought summaries only with extraBody: { generationConfig: { thinkingConfig: { includeThoughts: true } } }. Ollama models that only accept think: true or false need extraBody: { think: true } to turn thinking on.

Thinking appears as thinking blocks. Some providers sign their reasoning or return it encrypted and require it back unchanged in later requests, especially during tool loops. Nessi keeps that data on the blocks (signature, redacted, and details for OpenRouter's reasoning_details) and records the producing provider in message.provider; each provider sends it back only to itself. This covers Anthropic thinking, Gemini thought signatures, Mistral thinking chunks and OpenRouter reasoning items. Store assistant messages as they are, including these fields, and keep history append-only.

disableReasoning is deprecated. It keeps its original behavior: OpenAI-compatible providers send reasoning_effort: "low", Gemini sets a zero thinking budget and other providers ignore it. It is ignored when the call sets reasoningEffort.

Extra request fields

extraBody adds fields to the provider's request body for parameters Nessi does not model. Set it on the provider or per call; call values win. Plain objects merge deeply into what Nessi sends, other values replace it:

const provider = vllm("Qwen/Qwen3-32B", {
  baseURL: "http://localhost:8000/v1",
  extraBody: { chat_template_kwargs: { enable_thinking: false } },
});

Audio transcription

Use openAICompatibleTranscription() to upload an audio file to a service that implements the OpenAI-compatible /audio/transcriptions endpoint. Configure the service URL, API key and transcription model explicitly:

import { openAICompatibleTranscription } from "@k2b/nessi/ai";

const speech = openAICompatibleTranscription("whisper-large-v3", {
  baseURL: "https://api.scaleway.ai/v1",
  apiKey: process.env.SCW_SECRET_KEY,
});

const result = await speech.transcribe({
  file: Bun.file("./aufnahme.mp3"),
  filename: "aufnahme.mp3",
  language: "de",
  signal: AbortSignal.timeout(120_000),
});

console.log(result.text);

For OpenAI, set baseURL: "https://api.openai.com/v1", use your OpenAI key and a supported transcription model such as whisper-1. Local compatible services can omit apiKey. No environment variable is read automatically. Optional headers support gateways; apiKey overrides their Authorization header.

file accepts a Blob, File or Bun.file(). Use filename to supply an extension for an unnamed Blob or override the filename. Automatically derived filenames omit local directories. Omit language for automatic detection. Optional prompt supplies vocabulary or context when supported by the model. File formats, size limits and optional parameter support depend on the service; Nessi does not convert or split audio.

The result is { text: string }, including an empty string for an empty transcript. HTTP, connection and malformed-response errors reject the promise. Pass signal for cancellation or a timeout; cancellation preserves the signal's reason. There are no automatic retries or default timeout.

Transcription uses its own TranscriptionProvider contract with name, model and transcribe(request). Custom adapters can implement that interface for other protocols. It is separate from the chat provider passed to nessi(); pass the resulting text into the agent when needed. Timestamps and speaker identification are not exposed; for live audio see below.

The example follows Scaleway's audio API documentation. Keep API keys on the server when integrating a browser application.

Live transcription

vllmRealtimeTranscription() transcribes audio while it is still being recorded. It streams audio to vLLM's /v1/realtime WebSocket and yields text as the model recognizes it, for example with mistralai/Voxtral-Mini-4B-Realtime-2602:

import { vllmRealtimeTranscription } from "@k2b/nessi/ai";

const speech = vllmRealtimeTranscription("voxtral-realtime", {
  baseURL: "https://vllm.example.com/v1",
  apiKey: process.env.VLLM_API_KEY,
});

let text = "";
for await (const event of speech.stream({ audio: microphoneChunks, signal })) {
  if (event.type === "delta") text += event.text;
  if (event.type === "done") console.log(event.text, event.audioMs, event.usage);
}

audio is an async iterable of mono 16-bit little-endian PCM chunks (Int16Array or Uint8Array) at 16 kHz (speech.sampleRate). Send chunks as they are recorded; the transcript is finished when the iterable ends. delta events carry new text to append, and the final done event carries the complete transcript, the duration of the sent audio and token usage. Nessi does not resample audio. In a browser, an AudioContext created with { sampleRate: 16000 } records at the right rate.

The adapter uses the standard global WebSocket, so it runs in browsers, Bun, Deno and Node 22 or later. Sending apiKey or headers requires a WebSocket that accepts headers, which Bun does and browsers do not. Pass webSocket: (url, headers) => ... to supply your own WebSocket, for example from the ws package. Keep API keys on the server and relay browser audio through it.

Errors, a closed connection and a rejected handshake reject the iteration; cancellation through signal uses the signal's reason and closes the connection immediately. Stopping the iteration early closes it too. Nessi asks the audio iterable to stop, but cannot interrupt a pending read, so stop the microphone with the same signal. There are no automatic retries or reconnects.

Decision models

Decision models such as Cloudflare's Clef or Kev answer typed questions about a state with a probability for every option, in one forward pass and without generating text. They are fast and bounded, which makes them a good fit for routing, triage and checks before an LLM acts. Nessi speaks the System One API (POST /v1/systemone) that these models share:

import { systemOneDecision } from "@k2b/nessi/ai";

const router = systemOneDecision("clef-flash", {
  baseURL: "https://decide.example.com/v1",
  apiKey: process.env.DECISION_API_KEY,
});

const { answers } = await router.decide({
  state: "Checkout has been failing for every customer for the last hour.",
  questions: {
    urgent: { type: "noul", instructions: "Is this support request urgent?" },
    team: {
      type: "choice",
      instructions: "Which team should handle this request?",
      criteria: { billing: "Payments and refunds", technical: "Outages and errors", sales: "Plans" },
    },
    severity: { type: "score", instructions: "How severe is the impact?", criteria: ["None", "Minor", "Major", "Critical"] },
  },
});

answers.urgent.probability; // 0.91, `answers.urgent.value` is true at 0.5 or more
answers.team.choice;        // typed as "billing" | "technical" | "sales"
answers.team.confidence;    // probability of the chosen option
answers.severity.level;     // most likely level index, `score` is the weighted level

Question types:

  • noul: yes or no; optional criteria: { true, false } describe both answers. The answer has probability (of yes) and value.
  • choice: criteria maps option IDs to descriptions (or null). The answer has choice, confidence and probabilities per option.
  • score: criteria lists ordered levels, lowest first. The answer has the weighted score, the most likely level, confidence and probabilities per level index.

Option IDs are inferred from inline questions. For questions stored in a variable, keep the literal types with satisfies DecisionQuestions or as const. state accepts text or JSON-serializable data. Optional images ({ data: base64, mediaType }) are sent as data URLs; only multimodal models such as Clef accept them.

The result contains model, the typed answers, usage and raw, the unchanged provider response for fields Nessi does not map. Nessi checks that every question has a valid answer and otherwise rejects; it does not enforce provider limits such as the number of questions or options, which differ between models. HTTP, connection and malformed-response errors reject the promise, signal cancels with its reason, and there are no automatic retries.

For Clef on Cloudflare Workers AI, use the preset:

import { cloudflareDecision } from "@k2b/nessi/ai";

const router = cloudflareDecision("clef-flash", {
  accountId: process.env.CLOUDFLARE_ACCOUNT_ID!,
  apiToken: process.env.CLOUDFLARE_API_TOKEN!,
});

A decision model does not replace the agent loop; it decides before the loop acts. For example, pick the tools for a turn and fall back to all tools when the model is unsure:

const { answers } = await router.decide({
  state: userMessage,
  questions: {
    area: { type: "choice", instructions: "What is the request about?", criteria: { calendar: null, files: null, other: null } },
  },
});

const tools = answers.area.confidence >= 0.7 ? toolsByArea[answers.area.choice] : allTools;
const loop = nessi({ provider, store, systemPrompt, input: userMessage, tools });

Focused provider imports

import { anthropic } from "@k2b/nessi/ai/providers/anthropic";
import { openai } from "@k2b/nessi/ai/providers/openai";
import { openAICompatibleTranscription } from "@k2b/nessi/ai/providers/openai-compatible-transcription";
import { systemOneDecision } from "@k2b/nessi/ai/providers/systemone-decision";

Features

  • Turn-based agent loop with canonical block streaming events
  • Stable loopId correlation across all events from one agent loop
  • loop_start, turn_start, turn_end, and loop_end.aggregate for logical response grouping
  • Local loop.steer() and optional steering callbacks for steering at safe loop boundaries
  • Loop timing metadata for wall time, generation time, active tool time, action wait time, and output tokens/second
  • nessi.structured() for typed schema-valid task results
  • Server tools and client tools
  • Tool approval flow and explicit tool_action_request events
  • Tool execution start/end events with per-tool timeoutMs
  • Optional per-tool historical result representations for bounded future context
  • Structured issue events for provider errors, timeouts, malformed tool streams, and tool execution failures
  • Pluggable session store
  • Optional history compaction
  • Standalone compact() loop with loop_start, compaction_start, compaction_end, issue, and loop_end events
  • Optional token-credit budgeting
  • Provider adapters with shared complete() and stream() APIs
  • Audio transcription through configurable OpenAI-compatible services
  • Native adapters for OpenAI, OpenRouter, vLLM, Ollama, Anthropic, Mistral, and Gemini

Package layout

@k2b/nessi
  Agent loop, structured task helper, tools, stores, compaction, shared types

@k2b/nessi/ai
  Provider factories, provider types, complete(), stream(), responseFormat, transcribe()

@k2b/nessi/ai/providers/*
  Focused provider entrypoints