npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

llm-runtime

v0.8.0

Published

Runtime layer for application-owned LLM workflows with tool orchestration, MCP integration, and skill loading.

Readme

llm-runtime

llm-runtime is a TypeScript runtime layer for application-owned LLM workflows. It gives a host app one package boundary for provider calls, tool execution, MCP discovery, skills, and bounded agentic completion.

The package is intentionally not a full agent product. Your app still owns UI, persistence, permissions, transcript storage, workspace lifetime, and business policy. llm-runtime owns the provider/tool loop mechanics that should not be reimplemented in every harness.

For a human-oriented walkthrough of the codebase, start with the local project wiki: .wiki/index.md.

Installation

npm install llm-runtime

The package is ESM-only, requires Node.js 22+, and exposes a single root entrypoint.

Public Surface

The root entrypoint exports five runtime functions:

  • generate(...)
  • complete(...)
  • streamComplete(...)
  • createRuntime(...)
  • resolveToolCatalog(...)

It also exports the public type set for providers, messages, tools, runtime options, completion results, stream events, the effective tool catalog, MCP config, and provider config.

Lower-level loop internals, direct provider clients, recovery prompts, validation helpers, and tool-resolution helpers are not part of the root public API. If an app needs stable reusable dependencies and tool execution helpers, create a runtime and use the methods on that runtime instance.

Providers

Supported provider names:

  • openai
  • anthropic
  • google
  • azure
  • xai
  • openai-compatible
  • ollama

Provider configuration can be passed per call, through a provider map, or through a reusable runtime:

import { generate } from 'llm-runtime';

const response = await generate({
  provider: 'openai',
  model: 'gpt-5',
  providers: {
    openai: {
      apiKey: process.env.OPENAI_API_KEY!,
    },
  },
  messages: [
    { role: 'user', content: 'Summarize this in one paragraph.' },
  ],
});

console.log(response.content);

Runtime Model

Use createRuntime(...) when your app wants stable provider config, MCP config, skill roots, defaults, and registry state across many calls.

Put stable harness state in the runtime:

  • provider config
  • MCP config or MCP registry
  • skill roots or skill registry
  • default reasoningEffort
  • default toolPermission

Keep request-local state per call:

  • provider
  • model
  • messages
  • temperature
  • maxTokens
  • context.workingDirectory
  • context.abortSignal
  • webSearch
  • per-call builtIns, extraTools, or tools

Single-Turn Generation

generate(...) performs one provider call. It resolves the requested tool surface and passes it to the provider, but it does not execute returned tool calls or continue the conversation. Use it when the host wants to own the loop.

The returned LLMResponse is either:

  • type: 'text' with content
  • type: 'tool_calls' with tool_calls

Use complete(...) when the runtime should own repeated model calls, tool execution, and terminal control-tool handling.

In short: generate(...) asks the model once; complete(...) keeps working until the runtime reaches a completion, user-input, blocked, or bounded-stop condition.

Example:

import { createRuntime } from 'llm-runtime';

const runtime = createRuntime({
  providers: {
    openai: {
      apiKey: process.env.OPENAI_API_KEY!,
    },
  },
  skillRoots: ['/app/skills', '/workspace/.codex/skills'],
  defaults: {
    reasoningEffort: 'medium',
    toolPermission: 'auto',
  },
  mcpConfig: {
    servers: {
      docs: {
        command: 'node',
        args: ['docs-server.js'],
        transport: 'stdio',
      },
    },
  },
});

const response = await runtime.generate({
  provider: 'openai',
  model: 'gpt-5',
  messages: [
    { role: 'user', content: 'Read the project and identify the main runtime boundary.' },
  ],
  context: {
    workingDirectory: process.cwd(),
  },
  builtIns: {
    read_file: true,
    list_files: true,
    search_files: true,
  },
});

console.log(response.content);

await runtime.dispose();

Completion

complete(...) owns a bounded model/tool loop. It retries weak non-progressing responses, executes known tools, injects the runtime completion contract, and terminates through internal control tools:

  • final_answer
  • blocked

Those control tools are runtime-reserved. Do not define app tools with those names.

The loop continues under these rules:

  • normal tool call: execute the tool, append the tool result, and call the model again
  • final_answer: stop with status: 'completed'
  • ask_user_input: stop with status: 'tool_calls' so the host can ask/resume
  • known custom tool without an executor: stop with status: 'tool_calls' so the host can run it and resume
  • configured executable-tool approval cancellation: stop with status: 'cancelled' before executing the batch
  • blocked: stop with status: 'failed'
  • plain narration or intent text: keep going; narration is not completion
  • empty text: retry according to emptyTextRetryLimit
  • missing required action evidence: reject premature final text or final_answer and continue with recovery guidance
  • host mutating tool exposed: require a host mutating tool result before accepting final completion
  • repeated identical tool calls or maxIterations: stop with the corresponding bounded failure
  • host cancellation through context.abortSignal: abort the active model/tool path when the host decides the task should stop

builtIns only changes which package-owned tools are available. It does not decide whether the loop itself runs. builtIns: false with host-supplied extraTools or tools still gives complete(...) a valid loop: host tools remain executable, and the runtime still injects final_answer and blocked.

complete(...) returns:

  • status: 'completed' with output when the model reaches final_answer
  • status: 'tool_calls' when the host must handle a tool call, commonly required user input
  • status: 'cancelled' when configured executable-tool approval fails closed
  • status: 'failed' when the run is blocked, invalid, or otherwise cannot complete
  • status: 'max_iterations' when loop bounds stop the run

Example:

import { createRuntime } from 'llm-runtime';

const runtime = createRuntime({
  providers: {
    openai: {
      apiKey: process.env.OPENAI_API_KEY!,
    },
  },
});

const result = await runtime.complete({
  provider: 'openai',
  model: 'gpt-5',
  messages: [
    { role: 'user', content: 'Inspect the workspace and tell me where runtime completion is implemented.' },
  ],
  context: {
    workingDirectory: process.cwd(),
  },
  builtIns: {
    read_file: true,
    search_files: true,
    path_exists: true,
  },
  maxIterations: 12,
});

if (result.status === 'completed') {
  console.log(result.output);
}

if (result.status === 'tool_calls') {
  console.log(result.toolCalls);
}

streamComplete(...) runs the same completion path and yields lifecycle events. It emits model/tool events plus provider text, reasoning, tool-call argument, and final answer deltas when the provider adapter supplies them:

  • model_start
  • text_delta
  • reasoning_delta
  • tool_call_delta
  • answer_delta
  • assistant_message
  • tool_start
  • tool_result
  • tool_error
  • tool_calls
  • cancelled
  • completed
  • failed
  • raw

Do not concatenate text_delta and answer_delta into one user-visible message. They are different channels:

  • text_delta is raw assistant text emitted by the provider. In agentic runs that use the final_answer control tool, it can be draft or recovery text and should usually be treated as internal/debug output.
  • answer_delta is streamed from the final_answer control tool arguments. Hosts using the runtime's control protocol should display this as the final assistant answer.

For chat UIs that use final_answer, stream answer_delta to the visible assistant bubble. If no answer_delta arrives, use completed.result.output once as the fallback final text.

let streamedAnswer = '';

for await (const event of runtime.streamComplete({
  provider: 'openai',
  model: 'gpt-5',
  messages: [
    { role: 'user', content: 'Use tools if needed, then give the final answer.' },
  ],
  builtIns: {
    read_file: true,
    search_files: true,
    path_exists: true,
  },
})) {
  if (event.type === 'answer_delta') {
    streamedAnswer += event.delta;
    process.stdout.write(event.delta);
  }

  if (event.type === 'completed') {
    if (!streamedAnswer && event.result.output) {
      process.stdout.write(event.result.output);
    }
  }
}

Completed model boundary

Both complete(...) and streamComplete(...) accept an optional awaited onCompletedModelBoundary callback. The runtime invokes it after every normalized completed model iteration — tool-call output, control-tool (final_answer / blocked) output, rejected narration, and empty-response recovery included — and awaits it before another model call or any runtime-owned tool execution begins:

const result = await runtime.complete({
  provider: 'openai',
  model: 'gpt-5',
  messages: [{ role: 'user', content: 'Summarize the workspace.' }],
  onCompletedModelBoundary: async (boundary) => {
    // Durably commit boundary.iteration + boundary.response before the run
    // advances to the next model call or tool execution.
    await persistCompletedModelBoundary(boundary);
  },
});

The callback payload is CompletedModelBoundary:

  • iteration: 1-based model iteration index within the run.
  • requestMessages: the messages sent to the model for that iteration.
  • response: normalized operational data only — assistantMessage, type ("text" | "tool_calls"), toolCalls?, stopKind?, providerStopReason?, usage?, and warnings?. It never carries reasoning deltas, discarded draft text, raw secrets, or raw provider objects.

This is a durable-commit seam, not a persistence implementation: llm-runtime stores nothing. SQLite, conversation storage, checkpoints, event ledgers, and host policy all stay host-owned. Omit the callback and behavior is unchanged. Treat the payload as read-only: the runtime defensively copies the assistant message (including tool calls) and shallow-copies request messages, and host mutation of the payload is unsupported.

While a boundary callback is pending, the producer is stopped: iteration two and every tool executor wait for it. If the callback throws or rejects, the run stops before any later model turn or tool call with a distinguishable RuntimeCompletedModelBoundaryError (code === "completed_model_boundary_callback_failed", iteration, cause). Buffered complete(...) rejects with that error; streamComplete(...) emits one failed event whose result.errorCode is "completed_model_boundary_callback_failed" and whose iteration is the failing iteration. Do not consume the same streamComplete(...) stream from inside the callback, or the pending callback will deadlock the producer it is waiting on.

Tools

Tool sources are merged into one model-facing surface:

  • built-in runtime tools
  • app-provided extraTools
  • app-provided tools
  • MCP tools discovered from configured servers

Built-in tool names are reserved:

  • shell_cmd
  • load_skill
  • ask_user_input
  • web_fetch
  • read_file
  • write_file
  • list_files
  • search_files
  • create_directory
  • path_exists

Built-ins default to all package-owned tools for host convenience. Pass false to disable them, or pass a narrow map when the task should expose less:

  • omitting builtIns enables every built-in tool
  • builtIns: false enables no built-in tools
  • builtIns: true enables every built-in tool
  • pass an explicit per-tool map such as { read_file: true, search_files: true }
  • string shorthand modes such as builtIns: 'all' and builtIns: 'read-only' are not supported

Use small, task-specific maps:

const readOnlyBuiltIns = {
  load_skill: true,
  list_files: true,
  search_files: true,
  read_file: true,
  path_exists: true,
};

const writeFileBuiltIns = {
  ...readOnlyBuiltIns,
  create_directory: true,
  write_file: true,
};

const commandBuiltIns = {
  ...writeFileBuiltIns,
  shell_cmd: true,
};

Opt into write or command tools only when the task needs them. Do not use a broad preset for ordinary file inspection:

const result = await runtime.complete({
  provider: 'openai',
  model: 'gpt-5',
  messages: [
    { role: 'user', content: 'Run the project test command and summarize the result.' },
  ],
  context: {
    workingDirectory: process.cwd(),
  },
  builtIns: {
    read_file: true,
    search_files: true,
    path_exists: true,
    create_directory: true,
    write_file: true,
    shell_cmd: true,
  },
});

toolPermission: 'read' is a hard read-only boundary for package-owned mutating tools. It blocks write_file, create_directory, and shell_cmd even if those built-ins are exposed.

Prefer structured workspace tools over shell_cmd for routine file work:

  • list_files for directory listing
  • search_files for glob-like discovery
  • read_file for paginated file reads
  • path_exists for file or directory checks
  • create_directory for recursive directory creation when enabled

App tools can be passed as extraTools or tools. They are additive; they cannot override reserved built-ins or completion control tools.

const result = await runtime.complete({
  provider: 'openai',
  model: 'gpt-5',
  messages: [
    { role: 'user', content: 'Look up customer c_123 and summarize the account state.' },
  ],
  extraTools: [
    {
      name: 'lookup_customer',
      description: 'Look up a customer by id.',
      evidenceKind: 'read',
      parameters: {
        type: 'object',
        properties: {
          customerId: { type: 'string' },
        },
        required: ['customerId'],
        additionalProperties: false,
      },
      execute: async ({ customerId }) => {
        return { customerId, plan: 'enterprise', status: 'active' };
      },
    },
  ],
});

Runtime instances also expose resolveTools(...), executeToolCall(...), and executeToolCalls(...) for hosts that need to inspect or run the effective tool surface outside complete(...).

Effective Tool Catalog

Hosts that need to bind policy to the exact effective tool surface (startup handshake, authorization resume, schema hashing) should use the asynchronous catalog API. It resolves the same tool surface that completion and execution use — built-ins, custom extraTools/tools, and configured MCP tools under the same name resolution and precedence — and is available both as a runtime method and as a root-exported function:

import { createRuntime } from 'llm-runtime';

const runtime = createRuntime({
  mcpConfig: {
    servers: {
      search: {
        url: 'https://example.com/mcp',
        headers: {
          Authorization: `Bearer ${process.env.MCP_TOKEN}`,
        },
      },
    },
  },
});

const catalog = await runtime.resolveToolCatalog({
  builtIns: { read_file: true, search_files: true },
  extraTools: [{
    name: 'lookup_customer',
    description: 'Look up a customer by id.',
    evidenceKind: 'read',
    parameters: {
      type: 'object',
      properties: {
        customerId: { type: 'string' },
      },
      required: ['customerId'],
      additionalProperties: false,
    },
  }],
});

// catalog.tools maps effective tool name -> {
//   name, description, parameters (canonical model schema),
//   consequence ('none' | 'interaction' | 'read' | 'write' |
//     'external_action' | 'artifact'),
//   source ({ kind: 'builtin' | 'custom' | 'mcp', server?, tool? }),
//   digest (per-entry SHA-256), execute? (execution handle)
// }
// catalog.digest is the whole-catalog SHA-256.

Catalog guarantees:

  • Same surface as execution. The catalog is built by the same resolver used by complete(...), generate(...), stream(...), and executeToolCall(...), including MCP tools and MCP-overrides-custom precedence.
  • Deterministic model schemas. Entry parameters are a canonical (recursively key-sorted) form of the model-visible schema.
  • Resolved consequence class. consequence is the tool's declared evidenceKind when present; otherwise it follows the same name-based classification as completion, defaulting unknown tools (including MCP tools) to external_action.
  • Source identity. source.kind is builtin, custom, or mcp; MCP entries also expose source.server and source.tool (the original un-namespaced tool name). A custom MCP registry with a non-empty tool surface must implement resolveToolsWithSources() and provide both fields; resolveToolCatalog() rejects ambiguous provenance instead of producing a digest that could bind to the wrong server. Completion and execution retain the backward-compatible resolveTools() fallback.
  • Secret-free and digest-stable. Serialized catalog data never includes MCP server commands, arguments, environment variables, headers, URLs, credentials, secret references, or executor bodies. Digests are SHA-256 over the model-visible schema, consequence, and source identity only, so a digest change means a changed contract: a host can refuse to execute under a different definition after restart.

Deployment authorization, redaction, and secret-reference policy remain host-owned; the catalog deliberately does not carry host-policy fields.

Human Input

ask_user_input is the public human-input tool contract. It uses a structured questions[] payload:

{
  type?: "single-select" | "multiple-select";
  allowSkip?: boolean;
  questions: Array<{
    header: string;
    id: string;
    question: string;
    allowOther?: boolean;
    options: Array<{
      id: string;
      label: string;
      description?: string;
    }>;
  }>;
}

If you use a narrow builtIns map, include it when the model is allowed to ask the host for a human decision:

builtIns: {
  ask_user_input: true,
}

When completion needs host-owned user input, it returns status: 'tool_calls'. The host should surface the question, then resume by appending a normal tool-result message for the pending tool call and calling complete(...) again with the updated message list.

import {
  createAskUserInputResult,
  normalizeAskUserInputOutcome,
} from "llm-runtime";

const pending = {
  toolCallId: result.toolCalls![0].id,
  toolName: "ask_user_input",
  request: JSON.parse(result.toolCalls![0].function.arguments),
};
const outcome = normalizeAskUserInputOutcome(pending, {
  status: "answered",
  answers: {
    scope: "all",
  },
});

if (outcome.status === "cancelled") {
  // Stop this host workflow. Do not resume the model.
  return outcome;
}

const resumedMessages = [
  ...result.messages,
  createAskUserInputResult(pending, outcome),
];

The host owns rendering, waiting, timeout clocks, dismissal, and raw input collection. llm-runtime owns the request contract and normalization:

  • allowSkip: true means the host may dismiss the prompt. Dismissal is a cancelled outcome and never implies consent.
  • allowOther: true permits a non-empty free-form answer for that single-select question.
  • Multiple-select answers must contain one or more unique declared option IDs.
  • Invalid, partial, extra, skipped, dismissed, rejected, or timed-out responses normalize to status: "cancelled".

ask_user_input collects clarification and preferences. It does not authorize execution of a later tool call.

Host Execution Mode

Pass toolExecution: 'host' when the host (not the runtime) must execute every model tool call. In host mode complete(...) and streamComplete(...) parse and validate the complete known tool-call batch, execute none of it, and return the whole immutable ordered batch as status: 'tool_calls' for host execution:

const result = await runtime.complete({
  provider: 'openai',
  model: 'gpt-5',
  messages,
  toolExecution: 'host',
});

if (result.status === 'tool_calls') {
  // result.toolCalls is the complete, validated, immutable batch in order.
  // Execute each call in the host, then resume with normal tool-result messages.
}

Host-mode guarantees:

  • Complete known batch, unexecuted: every call must name a tool in the resolved tool set and carry arguments that parse and validate against the tool's declared schema. Validation is the package-owned schema validation (alias-normalization aware, matching execution-time validation), and the raw provider arguments are returned to the host as-is. If any call is unknown or invalid, the whole batch fails closed (per-call failure artifacts, tool_error events on the streaming path) and the loop continues so the model can correct — no partial batch is ever returned.
  • Mixed batches stay atomic: executable and host-owned (executor-free) calls in one batch are all returned together, in original order, with nothing executed — including calls whose definitions have executors.
  • Immutable batch: the returned toolCalls are frozen snapshots in the original order; the assistant tool-call message is included in messages.
  • No callbacks invoked: onToolApproval and onToolCall are not consulted in host mode because nothing executes. Control tools (final_answer, blocked) still terminate the loop as usual.
  • Resume without re-implementing the loop: append provider-valid role: 'tool' result messages for the pending calls and call complete(...) again (still with toolExecution: 'host' if the host keeps executing).

The default toolExecution: 'runtime' (or omitting the option) preserves the existing behavior: the runtime executes known executable tools inline, applies onToolApproval / onToolCall, and stops with status: 'tool_calls' only for known tools without executors. Host mode is a per-call execution policy: it introduces no session manager, run identity, or durable approval workflow.

Tool Approval

Use onToolApproval when the host must authorize executable tools. The callback receives the exact model-issued tool call and successfully parsed arguments:

const result = await runtime.complete({
  provider: "openai",
  model: "gpt-5",
  messages,
  onToolApproval: async ({ toolName, parsedArguments }) => {
    const decision = await showApprovalUI(toolName, parsedArguments);
    return decision === "approve"
      ? { decision: "approve" }
      : {
          decision: "cancel",
          reason: decision, // "rejected" | "dismissed" | "timeout"
        };
  },
});

if (result.status === "cancelled") {
  console.log(result.cancellation);
}

Approval is opt-in. When onToolApproval is omitted, executable tools retain automatic execution. When it is configured, only the exact { decision: "approve" } shape permits execution. Legacy booleans, { approved: true }, malformed values, callback errors, rejection, dismissal, and host-reported timeout all cancel without another model turn.

Approvals are collected for the complete executable tool-call batch before any tool in that batch runs. If one call is cancelled, none execute. The runtime does not start approval timers; a host-owned timeout must settle the callback with { decision: "cancel", reason: "timeout" }.

This is a breaking change from the pre-0.7 callback:

// Before
return { approved: true };

// 0.7+
return { decision: "approve" };

streamComplete(...) emits one cancelled terminal event for this outcome, not a failed event.

MCP And Skills

MCP servers are configured through mcpConfig. Both servers and legacy mcpServers shapes are accepted. URL-based servers default to streamable-http; stdio servers require a command.

const runtime = createRuntime({
  mcpConfig: {
    servers: {
      search: {
        url: 'https://example.com/mcp',
        headers: {
          Authorization: `Bearer ${process.env.MCP_TOKEN}`,
        },
      },
    },
  },
});

Skills are discovered from configured skillRoots and loaded through the load_skill built-in. Later skill roots have higher precedence when duplicate skill ids are found.

Skills add instruction context. They are not executable tools.

Web Search

Pass webSearch per call:

const response = await generate({
  provider: 'openai',
  model: 'gpt-5',
  providers: {
    openai: {
      apiKey: process.env.OPENAI_API_KEY!,
    },
  },
  messages: [
    { role: 'user', content: 'Use current public information to answer.' },
  ],
  webSearch: {
    searchContextSize: 'medium',
  },
});

Provider behavior:

  • openai, anthropic, and google receive provider-native web search options
  • azure, openai-compatible, xai, and ollama ignore unsupported web search on the current chat path and return a web_search_ignored warning
  • Gemini Google Search grounding is not combined with function calling; when both tools and webSearch are requested for google, tools win and web search is ignored with a warning
  • searchContextSize is forwarded for OpenAI-style requests and ignored by Anthropic and Gemini

Cleanup

Call runtime.dispose() when a runtime owns MCP clients:

const runtime = createRuntime({ mcpConfig });

try {
  await runtime.complete(request);
} finally {
  await runtime.dispose();
}

The host still owns temporary workspaces, transcript persistence, app-specific registries, and any resources it injected into the runtime.

Local Development

npm run build
npm run check
npm test

Useful scripts:

  • npm run build compiles src/ into dist/
  • npm run check runs TypeScript without emitting files
  • npm test runs the Vitest suite in tests/llm
  • npm run test:watch runs Vitest in watch mode
  • npm run test:e2e runs the live provider showcase
  • npm run test:e2e:dry-run validates showcase wiring without live provider calls
  • npm run test:e2e:azure runs Azure live-provider coverage
  • npm run test:e2e:azure:dry-run validates Azure showcase wiring
  • npm run test:e2e:gemini runs Gemini live-provider coverage
  • npm run test:e2e:gemini:dry-run validates Gemini showcase wiring
  • npm run test:e2e:turn-loop runs runtime completion showcase coverage
  • npm run test:e2e:turn-loop:dry-run validates turn-loop showcase wiring
  • npm run test:e2e:hardening runs deterministic hardening coverage without a live provider
  • npm run test:e2e:host-owned runs deterministic host-owned tool-call coverage without a live provider
  • npm run test:e2e:host-owned:gemini runs the host-owned tool-call coverage against Gemini 2.5 Flash by default

Live showcase runners expect a repo-local .env with the relevant provider credentials.