npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

agent-accelerator

v0.4.5

Published

High-performance SDK for AI agents with a unified API across AI providers. Optimized for high cache hit rates, observability, and seamless agent orchestration. Built to power production-grade agentic systems and harnesses for teams at any scale.

Readme

Agent Accelerator

npm version GitHub License: MIT

Stability notice: Agent Accelerator is pre-1.0 and under active development. Public APIs, provider transports, and configuration options may still change in breaking ways between releases. For production use, please pin your dependency to an exact version.

A thin, typed transport SDK for calling LLMs through a single Agent interface.

Agent Accelerator supports google, openrouter, openai, and any OpenAI-compatible endpoint using {PREFIX}_API_KEY and {PREFIX}_BASE_URL.

import { Agent } from "agent-accelerator";

const agent = new Agent({
  model: "google/gemini-3.5-flash-lite",
  instructions: "You are a concise research engineer.",
  cache: { retention: "short" },
});

const res = await agent.run("Explain HBM pricing in 3 bullets");

console.log(res.text);
console.log(res.usage);

Install

# Bun (recommended)
bun add agent-accelerator

# npm
npm install agent-accelerator

# pnpm
pnpm add agent-accelerator

# Yarn
yarn add agent-accelerator

Requires bun or node 22+.

Core dependencies are zod only. Document conversion additionally uses the optional @firecrawl/anydoc peer. All provider transports are native REST (fetch, no SDK).


Setup

GEMINI_API_KEY=...
OPENROUTER_API_KEY=...
OPENAI_API_KEY=...
# or OPENAI_BASE_API_KEY=...

MODEL="google/gemini-3.5-flash-lite"
SUB_AGENT_MODEL="google/gemini-3.5-flash-lite"

# Any OpenAI-compatible endpoint, no code change:
# MODEL="groq/llama-3.3-70b-versatile"
# GROQ_API_KEY="..."
# GROQ_BASE_URL="https://api.groq.com/openai/v1"

# Local, no key needed:
# MODEL="ollama/qwen2.5-coder"
# OLLAMA_BASE_URL="http://localhost:11434/v1"

MODEL and SUB_AGENT_MODEL are used when model or the dynamic worker model are omitted.

Explicit configuration always takes precedence over environment variables.


Quick Start

import { Agent } from "agent-accelerator";

const agent = new Agent({
  model: "google/gemini-3.5-flash-lite",
  instructions: "Be direct and technical.",
});

const res = await agent.run("Explain lock contention.");

console.log(res.text); // final answer
console.log(res.thinking); // reasoning trace, if any
console.log(res.usage); // input / output / cached / thinking / cost
console.log(res.raw.request.headers); // wire audit

Use instructions for the system prompt.

Keep instructions stable across turns to improve prefix-cache reuse.


Agent

new Agent(config: AgentConfig) is the main entry point.

An Agent holds :

  • instructions
  • tools
  • resolved model
  • thinking configuration
  • cache configuration
  • service tier
  • sessionId
  • conversation context

AgentConfig

| Field | Type | Meaning / Use Case | | ----------------- | ---------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | | model | string \| ModelSpec \| ModelProviderInstance | Model identifier, such as "google/gemini-3.5-flash-lite". Use ModelProvider.* or ModelSpec to pin a provider and credentials. Falls back to MODEL. | | instructions | string | System prompt. Keep it stable; place per-turn additions in additionalContext. | | tools | Record<string, ToolDefinition> \| ToolDefinition[] | Deterministic functions the model can call. See Tools. | | subagents | SubAgent[] \| Agent[] | Pre-defined workers. Each worker becomes a callable tool. Useful for fixed roles such as researcher or critic. | | subagentModel | string \| ModelSpec \| ModelProviderInstance | Developer-only model for dynamically spawned workers. Never choosable by the Main Agent. Overridden by dynamicSubagents.model. Falls back to SUB_AGENT_MODEL. | | dynamicSubagents | DynamicSubagentsConfig | Enables and constrains LLM-spawned stateless workers: { enabled, model?, maxSpawn?, thinkingLevel?, tools?, timeout? }. See Dynamic delegation. | | thinkingLevel | ThinkingLevel | none \| dynamic \| minimal \| low \| medium \| high \| xhigh \| max. Validated against the model catalog before a request. | | cache | CacheConfig | { retention, sessionId, cachedContentId, ttlSeconds }. Controls cache reuse. | | serviceTier | "flex" \| "priority" | Cost / priority routing where supported. Omit for standard routing. | | bypassInputFileModality | boolean | When true, file parts are converted client-side to Markdown for models lacking native support. Capable models still receive files natively. Defaults to false. | | maxTurns | number | Maximum model → tool → model loops per run. Infinite by default (0 or omitted = no limit). Hitting a finite limit with pending tool calls warns and sets finishReason: "max_turns". | | sessionId | string | Stable identifier used for cache affinity. Auto-generated when omitted. | | headers | Record<string,string> | Additional headers merged into every request. | | apiKey | string | Overrides environment-based API-key lookup for this agent. | | baseUrl | string | Overrides the default endpoint for this agent. | | stateless | boolean | When true, history is cleared before and after each run. Useful for one-shot evaluators. Defaults to false. | | persist | { dir?: string; file?: string } | Durable sessions: rewrites the session file after every step (user/assistant/tool/sub-agent). { dir } uses sessions/<sessionId>/session.json, { file } a single .session.json. Writes are atomic and never fail a turn. |

When dynamicSubagents.enabled is set without a resolvable worker model, construction throws SubAgentModelError.

Agent Methods

await agent.run(prompt, opts?) // non-streaming, or streaming when opts.stream is true
agent.ask(prompt, optsOrBool?) // string deltas when streaming, otherwise same as run
agent.stream(prompt, opts?) // AssistantMessageEventStream
agent.reset() // clears messages + signatures, keeps config
agent.exportSession(telemetry?) // storable snapshot for saveSessionFile
agent.importSession(saved) // restores messages + config from loadSessionFile
agent.track(trackingId) // live snapshot of one sub-agent worker (`NAME-{TrackingID}`)
agent.listTrackedSubAgents() // TrackingIDs of known workers
agent.subscribeToSubAgent(id, fn) // live worker updates, returns unsubscribe fn
agent.persistNow() // rewrites the persist target now (no-op unless persist is set)

prompt accepts string | ContentPart[].

Use ContentPart[] when sending images, audio, or video alongside text.

AgentRunOptions

| Field | Meaning / Use Case | | ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | stream | When true, returns a stream instead of a regular response promise. | | onDelta(delta, event) | Called for each text chunk. Use to render answers live. | | onThinkingDelta(delta, event) | Called for each reasoning chunk. Use to render reasoning separately. | | onEvent(event) | Called for every event: text_delta, thinking_delta, tool_call_complete, tool_result, subagent_complete, subagent_delta, usage, done, and error. | | wrapThinking | Wraps reasoning as <think>\n...\n</think>\n\n, allowing UIs to render it without maintaining separate state. | | onTurn(turn) | Called after every model/tool turn with { turns, text?, thinking?, toolCalls?, toolResults? }. Drives sub-agent step logs and persistence. | | additionalContext | Added only to the current user turn. Keeps the system prompt stable for better caching. | | sessionId | Per-run session override. | | headers | Per-run header merge. | | signal | AbortSignal used to cancel generation, tools, and subagents. |

Streaming

const res = await agent.run("Design an append-only log", {
  stream: true,
  wrapThinking: true,
  onThinkingDelta: (d) => process.stdout.write(d),
  onDelta: (d) => process.stdout.write(d),
  onEvent: (e) => {
    if (e.type === "subagent_complete") {
      console.log(`done: ${e.subagent!.name}`);
    }

    if (e.type === "tool_result") {
      console.log(`tool: ${e.toolResult!.name}`);
    }
  },
});

Manual Iteration

const stream = agent.stream("Hello");

for await (const e of stream) {
  if (e.type === "text_delta") {
    process.stdout.write(e.delta!);
  }

  if (e.type === "thinking_delta") {
    process.stdout.write(e.thinkingDelta!);
  }
}

const final = await stream.result();

stream.on("text_delta", (e) => console.log(e.delta));
stream.off("text_delta", handler);

SubAgent

SubAgent extends Agent and is designed for fixed worker roles.

import { Agent, SubAgent } from "agent-accelerator";

const researcher = new SubAgent({
  name: "researcher",
  instructions: "You collect quantifiable metrics.",
  model: "google/gemini-3.5-flash-lite",
});

const lead = new Agent({
  model: "google/gemini-3.7-flash",
  instructions: "Delegate research, then synthesize.",
  subagents: [researcher],
});

Differences from Agent

  • model resolves as config.model ?? SUB_AGENT_MODEL ?? MODEL and throws if empty.
  • cache defaults to { retention: "short" }.
  • .asTool(name?, desc?) / .toTool(name?, desc?) converts the subagent into a ToolDefinition for manual registration.
  • stateless: true is supported for one-shot judges that must not retain history.

SubAgentConfig is identical to AgentConfig.

Pre-defined vs Dynamic Delegation

Pre-defined delegation

Pass subagents: [a, b].

Each subagent is automatically registered as an internal tool on the parent that accepts:

{ task: string }

Dynamic delegation

subagents lists pre-defined workers (each becomes a callable tool). For LLM-spawned workers, configure dynamicSubagents:

const agent = new Agent({
  name: "Main Agent",
  model: process.env.MODEL,
  tools: { get_topic_brief },
  instructions: "You are the Main Agent",
  dynamicSubagents: {
    enabled: true,
    maxSpawn: 4,
    thinkingLevel: "low",
    tools: { get_weather, recent_news },
    timeout: 60000,
  },
});

A spawn_subagents tool is injected with the following task shape:

{
  tasks: {
    name,
    role?,
    instructions,
    task,
    tools?: string[],
    timeoutMs?: number
  }[]
}
  • maxSpawn — max workers per call. Extras are trimmed safely, and the limit is stated in the tool description plus the Main Agent's system prompt so the model knows it.
  • tools — developer-owned pool. Workers get zero tools unless the Main Agent grants a per-task tools subset by name (unknown names are ignored), keeping worker context small.
  • timeout (ms) — 0 = no limit, -1 = the Main Agent sets a per-task timeoutMs, >0 = fixed limit for every worker. Timed-out workers report an error entry; the rest of the batch still completes.
  • Workers are stateless: one task in, one result out, then shut down. No history, no recursion. The Main Agent cannot choose worker models or reasoning levels.
  • Worker names use UPPER-KEBAB (HBM-PRICING-ANALYST). Each worker gets a system-generated 32-char TrackingID, displayed as NAME-{TrackingID} (e.g. HBM-PRICING-ANALYST-9f2c…).
  • Workers stream internally: thinking/text deltas update the worker trace live (agent.track(id), subscribeToSubAgent) and surface on the parent stream as subagent_delta events.
const res = await agent.run("Audit auth pipeline and write a threat model");

console.log(res.subagents.map((s) => s.name));

Delegation Helpers

  • buildAgentTools(list) — wraps multiple agents and deduplicates names. Used internally by subagents.
  • createSubagentSpawnTool(parentAgent) — builds the dynamic subagent spawner. Rarely needed directly.
  • DynamicSubagentTask — { name, role?, instructions, task, tools?, timeoutMs? }, the shape produced for dynamic delegation. No model — workers run on the developer-configured model.

Tools

import { Agent, tool, z } from "agent-accelerator";

const agent = new Agent({
  model: "google/gemini-3.5-flash-lite",
  tools: {
    fetch_metrics: tool({
      description: "Fetch cluster metrics",
      input: z.object({
        clusterId: z.string(),
      }),
      execute: async ({ clusterId }) => ({
        clusterId,
        load: 0.42,
      }),
    }),
  },
});

Tool Helpers

  • tool({ name?, description, input?, parameters?, strict?, timeoutMs?, maxTries?, maxConcurrency?, execute }) — creates a ToolDefinition. Provide either a Zod input schema or raw JSON parameters. execute(input, ctx) may return arbitrary values; results are converted safely for model context. timeoutMs defaults to 0 (no time limit) and is a best-effort event-loop deadline; ctx.signal enables cooperative cancellation. maxTries accepts a number or numeric string; a positive value is the total attempt limit, while 0/omitted falls back to 3 attempts. maxConcurrency limits simultaneous calls for that tool (default pool: 8).
  • toStandardToolDeclarations(record | array) — converts tools to { name, description, parameters } for providers.
  • zodToJsonSchema(schema) — converts Zod schemas using native conversion with a fallback extractor.
  • cleanJsonSchema(schema) — removes $schema, $defs, and definitions, resolves $ref, and preserves explicit additionalProperties.
  • executeToolCalls({ tools, toolCalls, agentName?, parallel?, signal?, sessionId? }) — executes tool calls. Calls validate Zod input, use bounded parallelism, enforce per-tool timeouts, retry transient failures, recover namespace/camel-case tool-name aliases, and return ToolResultRecord[]. Missing tools and aborts are represented as error results rather than thrown.
  • convert_document_to_markdown — built-in tool that converts document files (PDF/Word/PowerPoint/Excel/OpenDocument/RTF/EPUB/CSV, local path or URL) to Markdown for models without native document parsing. Requires the optional @firecrawl/anydoc peer.

The agent loop also blocks an identical tool name and argument set when the model requests it in the immediately following turn. The synthetic error is returned to the model so it can reuse the prior result or change its arguments.

Tool Timing and Deadlines

Every ToolResultRecord includes durationMs, measured in whole milliseconds with a minimum displayed value of 1ms. The value covers the tool execution attempt (including retry/backoff time), but excludes queue wait and input-schema validation.

timeoutMs is specified in milliseconds:

const get_status = tool({
  name: "get_status",
  description: "Check user authentication status.",
  timeoutMs: 1,
  input: z.object({ username: z.string() }),
  execute: async ({ username }, ctx) => {
    // Pass ctx.signal to cancellable APIs such as fetch.
    return username === "Akshat Dwivedi" ? "Valid" : "Invalid";
  },
});

timeoutMs: 0 or an omitted value disables the deadline. Positive values are best-effort event-loop deadlines: JavaScript cannot forcibly stop arbitrary synchronous code, so asynchronous implementations should honor ctx.signal. A timeout is returned as an error tool result and is not retried indefinitely unless a positive maxTries is configured.

Tool timing is available after a run:

const response = await agent.run("Check status");
for (const result of response.toolResults) {
  console.log(`${result.name}: ${result.durationMs}ms`);
}

Tool Types

ToolExecutionContext:

{
  toolCallId,
  agentName?,
  signal?,
  sessionId?
}

ToolCallRecord:

{
  id,
  name,
  arguments,
  rawArguments?,
  thoughtSignature?,
  callId?,
}

ToolResultRecord:

{
  id,
  name,
  result,
  isError?,
  durationMs?
}

Providers

First-class provider IDs are:

  • google
  • openrouter
  • openai

Any other provider prefix is treated as an OpenAI-compatible custom provider.

Model Strings

  • "google/<id>" → Google AI Studio at generativelanguage.googleapis.com.
  • "openrouter/<scope>/<model>" or "scope/model:variant" → OpenRouter.
  • "openai/<id>" → OpenAI.
  • "groq/<id>", "ollama/<id>", etc. → custom providers using {PREFIX}_API_KEY and {PREFIX}_BASE_URL. The API key is optional for local endpoints.
  • A bare "some-model" uses OPENAI_BASE_URL when configured. Otherwise, catalog lookup is attempted, followed by an openrouter fallback.

Model resolution never silently selects a default model.

An empty model throws.

An unknown prefix with no catalog match still routes to a custom provider, allowing private model IDs and endpoints.

Registry

  • resolveModel("google/x" | ModelSpec) → { provider, modelId, modelSpec? }. Use to inspect routing before execution.
  • getProvider("groq") → returns a cached provider and automatically creates a custom provider for unknown prefixes.
  • ensureCustomProvider(prefix, { baseUrl?, apiKey?, name? }) → gets or creates a custom provider. Useful for multiple endpoints in one process.
  • normalizeProviderPrefix(s) → lowercases and trims a provider prefix.
  • ModelProvider.GoogleGenAI(model, apiKey?, { thinkingLevel?, baseUrl? }) — same shape is available for .OpenRouter, .OpenAI, .Custom, .OpenAICompatible, and .Generic.

Each returns a ModelProviderInstance:

{
  model,
  apiKey?,
  baseUrl?,
  thinkingLevel?
}

The result can be passed directly to Agent.model.

Use Custom for Groq, Together, Ollama, vLLM, and similar endpoints.

import { Agent, ModelProvider } from "agent-accelerator";

new Agent({
  model: ModelProvider.GoogleGenAI("gemini-3.5-flash-lite"),
});

new Agent({
  model: ModelProvider.Custom(
    "groq/llama-3.3-70b-versatile",
    process.env.GROQ_API_KEY,
    {
      baseUrl: process.env.GROQ_BASE_URL,
    },
  ),
});

Provider is the provider contract containing id, name, models, getModel, generate, and stream.

Its catalog-first getModel behavior includes a permissive fallback (createGenericModelSpec) so private model IDs never fail preflight.

Implement the Provider interface when adding a fully custom transport.

ProviderRequestOptions { apiKey?, baseUrl?, headers?, thinking?, cache?, serviceTier?, tools?, toolChoice?, signal?, sessionId?, env? } is the per-call options object passed to generate/stream. Agent builds it automatically.

Google

GoogleInteractionsProvider (aliased as GoogleAIStudioProvider) and GOOGLE_MODELS provide Google support over the Interactions API (POST {baseUrl}/interactions, streaming via ?alt=sse).

Turns chain statefully through previous_interaction_id per session, falling back to stateless full-history sends when the session switches providers or models.

Thinking levels map to thinking_level, with thinking_summaries enabled whenever thinking is active:

minimal | low | medium | high   (xhigh/max clamp to high; none is unsupported and omitted; dynamic omits)

Google-specific handling includes:

  • Isolating thoughtSignature per part for prefix stability.
  • Converting JSON Schema to OpenAPI 3.0 through stripSchemaForGoogle.
  • Implicit/automatic caching only — explicit retention and cachedContentId are warned about and dropped (the Interactions API defines no explicit cache primitives).

Helpers:

  • createExplicitCache({ model, systemInstruction?, contents?, tools?, displayName?, ttlSeconds?, expireTime?, apiKey?, baseUrl? })
  • isValidThoughtSignature(sig)
  • retainThoughtSignature(existing, incoming)
  • extractGoogleThoughtSignature(obj)
  • stripSchemaForGoogle(schema)
  • clearInteractionChains(sessionId?)

OpenAI

OpenAIResponsesProvider (aliased as OpenAIProvider) and OPENAI_MODELS provide OpenAI support over the Responses API (POST {baseUrl}/responses).

Every turn is stateless: the full history is sent explicitly with store: false, so no previous_response_id, background, or conversation chaining is ever used.

Thinking levels map to reasoning.effort verbatim:

none | minimal | low | medium | high | xhigh | max   (dynamic omits — server default)

It also:

  • passes service_tier (flex | priority);
  • pins cache affinity with prompt_cache_key (clamped to 64 chars);
  • rejects video parts up front with a one-line error (the Responses wire carries text/image/file only);
  • never sends temperature, token caps, or background/conversation fields.

OpenRouter

OpenRouterChatCompletionsProvider (aliased as OpenRouterProvider / OpenRouter) and OPENROUTER_MODELS provide OpenRouter support over Chat Completions (POST {baseUrl}/chat/completions).

Every turn is stateless: the full messages[] array is sent explicitly — no chaining primitives exist on this endpoint.

Session affinity is a top-level body session_id plus the x-session-id header fallback (body takes precedence per OpenRouter docs).

Thinking levels map to reasoning.effort verbatim:

none | minimal | low | medium | high | xhigh | max   (dynamic omits — server default)

It also:

  • passes service_tier (flex | priority);
  • enables the file-parser plugin when file parts are present (remote file URLs are additionally surfaced in text);
  • maps tool_choice (auto omitted; none, required, and function pins supported);
  • reports cached_tokens / cache_write_tokens / reasoning tokens, preferring provider-reported cost when present.

serviceTier values other than flex / priority are omitted (standard routing).

Custom

OpenAICompatibleProvider is also exposed through the CustomProvider, createCustomProvider, and createOpenAICompatibleProvider(prefix, opts?) aliases.

CustomProviderOptions:

{
  name?,
  baseUrl?,
  apiKey?,
  defaultBaseUrl?
}

createGenericModelSpec(provider, modelId) creates a permissive placeholder so private model IDs can pass validation and window checks.

Provider wire types are exported for inspecting raw requests and responses through:

res.raw.request.body
res.raw.response.body

Examples include:

  • OpenAIChatCompletionRequest
  • GoogleGenerateContentRequest
  • OpenRouterChatRequest

Thinking

Thinking uses a single thinkingLevel flag:

ThinkingLevel =
  | "none"
  | "dynamic"
  | "minimal"
  | "low"
  | "medium"
  | "high"
  | "xhigh"
  | "max";

Thinking Levels

  • none — disables thinking where supported.
  • dynamic — allows the model to decide.
  • minimal, low, medium, high, xhigh, max — request increasing levels of reasoning where supported.

Thinking can be configured on the Agent or through ModelProvider.*(..., { thinkingLevel }).

The per-run thinkingLevel option can override the agent default for one request.

Internally, thinking is normalized to:

ThinkingConfig {
  enabled?,
  level?,
  budgetTokens?,
  includeThoughts?
}

Catalog Helpers

  • getModelThinkingInfo(provider, modelId) → { supportsThinking, reasoningOptions?, allowedLevels, supportsDisable, description }. Useful for building UI selectors.
  • validateModelThinking(provider, modelId, level) → throws ThinkingLevelError { provider, modelId, requestedLevel, allowedLevels, supportsThinking } with a fix hint. Automatically called by run and stream. Catch it and use err.allowedLevels to offer valid options.
  • resolveEffectiveThinking(base, override?) → resolves a per-run thinkingLevel override into a ThinkingConfig without mutating agent config.
  • getModelsForProvider("google") → returns known model specifications.

Unknown models allow all thinking levels.

Fixed-reasoning models reject none.


Cache

CacheConfig {
  retention?,
  cachedContentId?,
  sessionId?,
  ttlSeconds?
}
CacheRetention =
  | "implicit"
  | "short"
  | "medium"
  | "long";

Retention

  • implicit — automatic prefix reuse without an explicit cache object or storage fee. This is the only retention the native adapters act on (by doing nothing special — stable prompts plus session affinity do the work).
  • short / medium / long — TTL hints (5m / 1h / 12h) consumed only by the opt-in createExplicitCache helper. Provider adapters warn about and drop any non-implicit retention: none of the native transports expose retention control.

Session Affinity

sessionId pins provider affinity using mechanisms such as:

  • x-session-id (+ x-client-request-id) headers — all providers, best effort
  • top-level body session_id — OpenRouter (takes precedence over the header)
  • prompt_cache_key — OpenAI only (custom endpoints get headers only; strict ones reject unknown body fields)

Reuse the same session ID across turns when cache affinity is desired.

In browser runtimes custom x-* headers are stripped to avoid CORS preflights: OpenAI keeps affinity through its body key, while custom endpoints lose affinity entirely on browsers.

Explicit Caches

createExplicitCache mints a Google cachedContents/... resource directly (TTL via retention/ttlSeconds).

cachedContentId is accepted and carried in context, but no native adapter currently sends it — explicit references are warned about and dropped, so stable prompts plus session affinity remain the cache-reuse path.

ttlSeconds overrides the retention mapping.

Provider Cache Mechanisms

  • Google keeps system prompts and tools stable and isolates signatures (implicit caching only).
  • OpenAI pins affinity with prompt_cache_key.
  • OpenRouter pins affinity with body session_id plus headers (sticky routing from the first request).
  • Custom endpoints get headers-only affinity; anything else cache-related is warned about and dropped.

The following cache helpers are exported:

  • getPromptCacheRetention
  • retentionToTtlSeconds
  • clampCacheKey

For the best cache reuse, keep instructions and the tool set stable. Put changing data in user messages or additionalContext.


serviceTier

serviceTier = "flex" | "priority";

Omit serviceTier for standard routing.

Where supported, it is passed as service_tier or translated into provider-specific routing.

  • flex — suitable for batch-tolerant workloads.
  • priority — suitable for latency-sensitive workloads.

Streaming and Events

AssistantMessageEventStream implements AsyncIterable<StreamEvent>.

It provides:

.result(): Promise<AgentResponse>
.push(...)
.end(...)
.fail(...)
.on(...)
.off(...)
.cancel()
.onCancel(...)
.isCancelled()

StreamEvent

StreamEvent {
  type,
  delta?,
  thinkingDelta?,
  toolCall?,
  toolResult?,
  subagent?,
  subagentTrackingId?,
  usage?,
  responseId?,
  finishReason?,
  error?,
  partialText?,
  partialThinking?,
  raw?
}

Event types:

start
text_start
text_delta
text_end
thinking_start
thinking_delta
thinking_end
tool_call_start
tool_call_delta
tool_call_complete
tool_result
subagent_complete
subagent_delta
usage
done
error

Emitted in practice: start, text_delta, thinking_delta, tool_call_complete, tool_result, subagent_complete, subagent_delta, usage, done, error. The rest exist in the type for forward compatibility.

Use:

  • text_delta for answer output.
  • thinking_delta for reasoning output.
  • tool_call_complete / tool_result for tool progress.
  • subagent_complete for each completed worker during streaming multi-agent runs.
  • subagent_delta for live worker thinking/text (subagentTrackingId + delta/thinkingDelta + partials). Render per-worker; never merge into the parent answer.
  • usage for interim usage counts.
  • done for final usage and completion information.

Breaking out of for await auto-cancels. Cancellation joins your AbortSignal and stream cancellation into one linked controller, so controller.abort() and stream.cancel() both stop the HTTP request. Abort rejects with AbortError (never retried) and never resolves partial results. agent.run(prompt, { stream: true, signal }) exposes the same cancel handle.

SSEParser provides:

feed(chunk)
flush()

It parses provider SSE streams and is only needed when implementing custom transports.


Response, Usage, and Raw Data

AgentResponse

AgentResponse {
  text,
  thinking?,
  thoughtSignature?,
  toolCalls,
  toolResults,
  subagents,
  usage,
  responseId?,
  model,
  provider,
  finishReason?,
  durationMs,
  raw,
  turns
}

AgentResponse also provides:

.toString()
.toJSON()

SubAgentExecutionMetadata

SubAgentExecutionMetadata {
  name,
  role?,
  task,
  model,
  provider,
  durationMs,
  usage,
  turns,
  finishReason?,
  responseId?,
  text,
  thinking?,
  toolCalls?,
  raw?,
  isError?,
  error?,
  trackingId?,
  sessionId?,
  parentSessionId?,
  steps?
}

finishReason is "max_turns" when a finite maxTurns limit stopped the run with tool calls still pending (plus a console warning naming the dropped tools).

SubAgentStep

SubAgentStep {
  turn,
  type, // "assistant" | "tool_call" | "tool_result"
  name?,
  text?,
  thinking?,
  isError?,
  timestamp,
  partial? // true while the worker is still streaming this step
}

agent.track(id) returns these live; partial entries finalize when the worker's turn completes.

TokenUsage

TokenUsage {
  inputTokens,
  outputTokens,
  totalTokens,
  cachedTokens?,
  cacheReadTokens?,
  cacheWriteTokens?,
  thinkingTokens?,
  cost?: {
    inputCost?,
    outputCost?,
    cacheReadCost?,
    cacheWriteCost?,
    totalCost?
  }
}

Cost calculation prefers provider-reported totals. When unavailable, costs are computed using catalog pricing.

Subagent usage is rolled into the parent totals.

ProviderRawData

ProviderRawData {
  request: {
    url,
    method,
    headers,
    body
  },
  response?: {
    status,
    statusText,
    headers,
    body
  }
}

Sensitive keys are redacted.


Errors

Provider failures throw AgentAccelProviderError — one actionable line, never a wire dump:

[openrouter/nvidia/nemotron-3.5-lightning:free] request failed (404): No endpoints found that support input video

Raw details (statusCode, url, truncated body) stay attached as non-enumerable properties: available programmatically, invisible in runtime dumps. URL query strings are stripped (keys sometimes live there); headers and cookies are never attached. Helpers:

  • toConciseProviderError(err, providerId, modelId) — collapse any provider error.
  • assertModalitiesSupported(context, providerId, modelId) — pre-request modality gate.

Examples print failures as exactly one line via examples/_shared.ts:

const res = await agent.run([...]).catch(fail);
// ✖ [provider/model] request failed (404): ...
// exit 1, no stack dump

Model Catalog & Dynamic 12-Hour TTL Sync

Agent Accelerator features an adaptive, dynamically synchronized model catalog powered by models.dev. To avoid shipping a bloated 4.4 MB static JSON file with production builds, the SDK employs a high-performance 12-hour TTL local caching architecture:

  • Automated 12-Hour Cache Validation: When .run(), .ask(), or .stream() executes, the runtime checks src/data/models.dev.json. If the cache timestamp is within 12 hours, it reads from disk with zero network delay. When the TTL expires, it transparently synchronizes with https://models.dev/api.json.
  • Git & Package Safety: The dynamic cache file (src/data/models.dev.json) is gitignored and excluded from production packages.

Developer Catalog Controls

import {
  refreshModelCatalog,
  getCatalogStatus,
  setCatalogTTL,
  getModelFromCatalog,
  getModelThinkingInfo,
} from "agent-accelerator";

// Force an immediate refresh from models.dev:
await refreshModelCatalog({ force: true });

// Customize default TTL (e.g. 24 hours):
setCatalogTTL(24 * 60 * 60 * 1000);

// Inspect cache health and metadata:
const status = getCatalogStatus();
console.log(status.modelCount, status.providerCount, status.isExpired);

CLI Refresh

bun run update-models              # Refresh model catalog (force or expired)
bun src/update-models.ts --force  # Force re-download
bun src/update-models.ts --ttl=24h # Refresh with custom TTL

Upstream data lags on some inputs (e.g. gpt-4o accepts audio). Record verified corrections in MODALITY_OVERRIDES in src/models/catalog.ts (keyed provider/model, merged over the snapshot) — never edit the JSON directly, a refresh would wipe it.

Catalog Helpers

  • getModelFromCatalog(provider, modelId) → ModelSpec | undefined, including alias and global fallback handling. Use it for context-window, pricing, and modality checks.
  • getModelsForProvider(provider) → returns ModelSpec[].
  • getModelThinkingInfo(provider, modelId) → inspect permitted thinking levels.
  • refreshModelCatalog(options?) → programmatically sync with upstream models.dev.
  • getCatalogStatus() → diagnostic info on active catalog, provider count, and TTL expiry.

ModelSpec

ModelSpec {
  id,
  provider,
  name,
  description?,
  family?,
  api?,
  contextWindow,
  maxOutputTokens,
  limit?,
  cost?,
  modalities?,
  reasoning?,
  reasoning_options?,
  tool_call?,
  capabilities,
  pricing?,
  raw?
}

contextWindow and maxOutputTokens mirror:

limit.context
limit.output

ModelCapabilities

ModelCapabilities {
  supportsThinking?,
  supportsThinkingBudget?,
  supportsThinkingLevel?,
  supportsImplicitCaching?,
  supportsExplicitCaching?,
  supportsLongCacheRetention?,
  supportsParallelToolCalls?,
  supportsStreaming?,
  modalities?,
  supportsReasoningToggle?,
  supportsReasoningEffort?
}

Messages and Media

Message

Message {
  role: "system" | "user" | "assistant" | "tool",
  content: string | ContentPart[],
  name?,
  thoughtSignature?
}

ProviderContext

ProviderContext {
  systemPrompt?,
  messages,
  tools?,
  cachedContentId?
}

ContentPart

Supported content parts include:

  • TextPart { type:"text", text, thoughtSignature? } — plain text.
  • ThinkingPart { type:"thinking", thinking, thoughtSignature? } — preserved reasoning.
  • ToolCallPart { type:"tool_call", id, name, arguments, rawArguments?, thoughtSignature? }
  • ToolResultPart { type:"tool_result", id, name, result, isError? }
  • ImagePart { type:"image", image, mimeType? }
  • AudioPart { type:"audio", audio, mimeType? }
  • VideoPart { type:"video", video, mimeType? }
  • FilePart { type:"file", file, mimeType?, filename? } — PDFs/documents.

image, audio, video, and file inputs accept:

  • data URLs
  • remote http(s) URLs
  • local paths
  • raw base64
  • Uint8Array
  • ArrayBuffer

normalizeMediaInput(input, mime?) returns:

{
  mimeType,
  base64Data,
  dataUrl
}

inferMimeType(path) infers the MIME type from a file extension.

Media normalization is handled automatically by providers (mapped to provider-native media shapes). Thinking traces are echoed in follow-up turns only where accepted — strict endpoints receive tool calls without reasoning_content.

Provider matrix: text + image + wav/mp3 audio + PDF work on all supported providers. Video works on Gemini only.

Before any network call, the executor checks image / audio / video / file-as-PDF parts against the catalog's modalities.input for that provider/model and fails fast with a one-line error naming the gap (assertModalitiesSupported). Models absent from the catalog (custom providers, dynamic routers) are skipped — the provider endpoint decides and its verdict surfaces as a concise error (see Errors). Verified catalog corrections live in MODALITY_OVERRIDES (src/models/catalog.ts), never in the gitignored snapshot.


Utils

  • createSessionId(prefix="accel") — creates a UUID-based session ID clamped to 64 characters for affinity.
  • createTrackingId() — system-generated 32-char TrackingID for one sub-agent worker (NAME-{TrackingID}).
  • createChildSessionId(parentId, tag) / createFixedChildSessionId(parentId, name) / isSessionDescendant(child, parent) — provider-safe (≤64 chars) child session IDs with parent-hash lineage.
  • getEnv(key, fallback?) — resolves an environment variable.
  • getApiKey(provider, explicit?, env?) — resolves API keys using the configured environment lookup order.
  • buildSessionHeaders(provider, cache?, custom?, sessionId?) — builds provider-specific affinity headers. Normally handled automatically.
  • withRetries(fn, { maxRetries?, maxRetryDelayMs?, signal?, label? }) / isTransientError(err) — bounded retries for transient provider failures (defaults: 2 retries, 5s cap, aborts never retried).
  • convertDocumentToMarkdown(input, { filename?, format?, mimeType?, maxChars? }) — converts PDF/Word/PowerPoint/Excel/OpenDocument/RTF/EPUB/CSV to Markdown via the optional @firecrawl/anydoc peer. Throws DocumentConversionError on failure.
  • convert_document_to_markdown — built-in model-callable version of the above (returns a <Document> envelope); auto-registered with bypassInputFileModality: true.
  • z — re-exported from Zod so tools do not require a separate Zod import.

Examples

bun run examples/05-chat.ts
# persistent CLI, /model "..." /level <lvl> /help /exit

bun run examples/04-sub-agents.ts
# fixed researcher + critic pipeline

bun run examples/03-multi_agent.ts
# dynamic spawn_subagents demo

bun run examples/01-metadata.ts
# usage + raw inspection

bun run examples/02-function_calling.ts
# single tool call

bun run examples/06-multimodal_image.ts [./photo.png]
# image input (remote URL default, local path optional)

bun run examples/07-multimodal_audio.ts
# audio input

bun run examples/09-multimodal_document.ts
# PDF/file input

bun run examples/08-multimodal_video.ts
# video input (video-capable model required)

bun run examples/10-document_markdown.ts
# document → Markdown preprocessing (any model, optional @firecrawl/anydoc peer)

bun run examples/11-session-identity.ts
# sub-agent TrackingIDs, live worker boxes, lineage, track(), durable sessions

Every example ends its run() with .catch(fail) (examples/_shared.ts), so failures print one line and exit 1 — no stack dumps.

The chat example persists conversations to:

sessions/<sessionId>/session.json  (+ media/ for images, audio, video, files)

It resumes the most recent session on next launch (legacy .session.json/.session.jsonl files still load) and prints a resume banner:

↺ Previous session loaded • accel-1a2b… (.session.json) • 12 messages • google/gemini-3.5-flash-lite • total-CH82.4% • $0.013 total

It also displays per-turn and session totals on one line:

↑turn ↓turn CRturn turn-CH% | ↑total ↓total CRtotal total-CH% $total(+turn) ctx%/window

Session persistence

import {
  Agent,
  loadSessionFile,
  saveSessionFile,
  SessionTelemetry,
  formatSessionBanner,
} from "agent-accelerator";

const saved = loadSessionFile(".session.json");
const telemetry = SessionTelemetry.fromSaved(saved);

const agent = new Agent({
  model: saved?.model ?? process.env.MODEL,
  sessionId: saved?.sessionId,
});

if (saved) agent.importSession(saved);

const res = await agent.run("Hello");
telemetry.add(res.usage, agent.modelStringOrSpec as string);
saveSessionFile(".session.json", agent, telemetry);
  • agent.exportSession(telemetry) / agent.importSession(saved) round-trip messages, system prompt, model, thinking level, instructions, cache, worker model, and sub-agent traces without touching AgentContext internals.
  • SessionTelemetry accumulates input / output / cacheRead / cacheWrite / reasoning / cost, with clamped turnHitRate() / totalHitRate(), a dual formatBar(res), and formatSessionBanner(saved, telemetry, path) for startup.
  • Files are pretty-printed JSON (2-space indent), written atomically (temp + rename). Cost prefers provider-reported totals and falls back to catalog pricing.
  • Binary media (bytes, data URLs, base64) is extracted to media/ on save via saveSessionDir("sessions", agent, telemetry) → sessions/<sessionId>/{session.json, media/*}; remote URLs and local paths stay references. loadSessionDir(dir) resolves media/… refs back to absolute paths, and findLatestSessionDir("sessions") resumes the most recent session.
  • Sub-agent traces persist in a dedicated subagents section keyed by TrackingID ({ trackingId, name, sessionId, parentSessionId, status, task, turns, usage, steps, text }). Files stay lean: redundant rawArguments and repeated step text are dropped on save (rebuilt/kept live in memory).
  • With persist: { dir } or { file }, the session file is rewritten after every step, so a crash loses at most the in-flight step.

Scripts and Structure

Scripts

bun run typecheck     # tsc --noEmit
bun test              # bun test test/
bun run update-models # refresh model catalog cache (supports --force, --ttl=24h)

Project Structure

src/
├── agent/      # Agent, context, loop, delegation, subagent
├── session/    # Session persistence + telemetry (snapshots, hit rates, .session.json store)
├── providers/  # native REST adapters (google/openai/openrouter/openai-compat) + registry + canonical contract
├── models/     # Dynamic catalog cache, parser, verified overrides
├── data/       # Dynamic model catalog cache (gitignored, excluded from bundle)
├── tools/      # tool(), schema, executor
├── streaming/  # event stream, SSE parser
├── types/      # agent, core, message, model, response, tool
└── utils/      # base64, cache, env, headers, media, serialization, session, thought-signature, documents, retry, errors
examples/
├── 01-metadata.ts  02-function_calling.ts  03-multi_agent.ts
├── 04-sub-agents.ts  05-chat.ts  06-multimodal_image.ts
├── 07-multimodal_audio.ts  08-multimodal_video.ts
├── 09-multimodal_document.ts  10-document_markdown.ts
├── files/  prompts/  research-agent.ts  _shared.ts
test/            # unit + mocked-provider tests mirroring src/

License

MIT License — see LICENSE for details.

Copyright (c) 2026 Sashvat Bharat.