agent-accelerator
v0.4.5
Published
High-performance SDK for AI agents with a unified API across AI providers. Optimized for high cache hit rates, observability, and seamless agent orchestration. Built to power production-grade agentic systems and harnesses for teams at any scale.
Maintainers
Readme
Agent Accelerator
Stability notice: Agent Accelerator is pre-1.0 and under active development. Public APIs, provider transports, and configuration options may still change in breaking ways between releases. For production use, please pin your dependency to an exact version.
A thin, typed transport SDK for calling LLMs through a single Agent interface.
Agent Accelerator supports google, openrouter, openai, and any OpenAI-compatible endpoint using {PREFIX}_API_KEY and {PREFIX}_BASE_URL.
import { Agent } from "agent-accelerator";
const agent = new Agent({
model: "google/gemini-3.5-flash-lite",
instructions: "You are a concise research engineer.",
cache: { retention: "short" },
});
const res = await agent.run("Explain HBM pricing in 3 bullets");
console.log(res.text);
console.log(res.usage);Install
# Bun (recommended)
bun add agent-accelerator
# npm
npm install agent-accelerator
# pnpm
pnpm add agent-accelerator
# Yarn
yarn add agent-acceleratorRequires bun or node 22+.
Core dependencies are zod only. Document conversion additionally uses the optional @firecrawl/anydoc peer. All provider transports are native REST (fetch, no SDK).
Setup
GEMINI_API_KEY=...
OPENROUTER_API_KEY=...
OPENAI_API_KEY=...
# or OPENAI_BASE_API_KEY=...
MODEL="google/gemini-3.5-flash-lite"
SUB_AGENT_MODEL="google/gemini-3.5-flash-lite"
# Any OpenAI-compatible endpoint, no code change:
# MODEL="groq/llama-3.3-70b-versatile"
# GROQ_API_KEY="..."
# GROQ_BASE_URL="https://api.groq.com/openai/v1"
# Local, no key needed:
# MODEL="ollama/qwen2.5-coder"
# OLLAMA_BASE_URL="http://localhost:11434/v1"MODEL and SUB_AGENT_MODEL are used when model or the dynamic worker model are omitted.
Explicit configuration always takes precedence over environment variables.
Quick Start
import { Agent } from "agent-accelerator";
const agent = new Agent({
model: "google/gemini-3.5-flash-lite",
instructions: "Be direct and technical.",
});
const res = await agent.run("Explain lock contention.");
console.log(res.text); // final answer
console.log(res.thinking); // reasoning trace, if any
console.log(res.usage); // input / output / cached / thinking / cost
console.log(res.raw.request.headers); // wire auditUse instructions for the system prompt.
Keep instructions stable across turns to improve prefix-cache reuse.
Agent
new Agent(config: AgentConfig) is the main entry point.
An Agent holds :
instructionstools- resolved model
- thinking configuration
- cache configuration
- service tier
sessionId- conversation
context
AgentConfig
| Field | Type | Meaning / Use Case |
| ----------------- | ---------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- |
| model | string \| ModelSpec \| ModelProviderInstance | Model identifier, such as "google/gemini-3.5-flash-lite". Use ModelProvider.* or ModelSpec to pin a provider and credentials. Falls back to MODEL. |
| instructions | string | System prompt. Keep it stable; place per-turn additions in additionalContext. |
| tools | Record<string, ToolDefinition> \| ToolDefinition[] | Deterministic functions the model can call. See Tools. |
| subagents | SubAgent[] \| Agent[] | Pre-defined workers. Each worker becomes a callable tool. Useful for fixed roles such as researcher or critic. |
| subagentModel | string \| ModelSpec \| ModelProviderInstance | Developer-only model for dynamically spawned workers. Never choosable by the Main Agent. Overridden by dynamicSubagents.model. Falls back to SUB_AGENT_MODEL. |
| dynamicSubagents | DynamicSubagentsConfig | Enables and constrains LLM-spawned stateless workers: { enabled, model?, maxSpawn?, thinkingLevel?, tools?, timeout? }. See Dynamic delegation. |
| thinkingLevel | ThinkingLevel | none \| dynamic \| minimal \| low \| medium \| high \| xhigh \| max. Validated against the model catalog before a request. |
| cache | CacheConfig | { retention, sessionId, cachedContentId, ttlSeconds }. Controls cache reuse. |
| serviceTier | "flex" \| "priority" | Cost / priority routing where supported. Omit for standard routing. |
| bypassInputFileModality | boolean | When true, file parts are converted client-side to Markdown for models lacking native support. Capable models still receive files natively. Defaults to false. |
| maxTurns | number | Maximum model → tool → model loops per run. Infinite by default (0 or omitted = no limit). Hitting a finite limit with pending tool calls warns and sets finishReason: "max_turns". |
| sessionId | string | Stable identifier used for cache affinity. Auto-generated when omitted. |
| headers | Record<string,string> | Additional headers merged into every request. |
| apiKey | string | Overrides environment-based API-key lookup for this agent. |
| baseUrl | string | Overrides the default endpoint for this agent. |
| stateless | boolean | When true, history is cleared before and after each run. Useful for one-shot evaluators. Defaults to false. |
| persist | { dir?: string; file?: string } | Durable sessions: rewrites the session file after every step (user/assistant/tool/sub-agent). { dir } uses sessions/<sessionId>/session.json, { file } a single .session.json. Writes are atomic and never fail a turn. |
When dynamicSubagents.enabled is set without a resolvable worker model, construction throws SubAgentModelError.
Agent Methods
await agent.run(prompt, opts?) // non-streaming, or streaming when opts.stream is true
agent.ask(prompt, optsOrBool?) // string deltas when streaming, otherwise same as run
agent.stream(prompt, opts?) // AssistantMessageEventStream
agent.reset() // clears messages + signatures, keeps config
agent.exportSession(telemetry?) // storable snapshot for saveSessionFile
agent.importSession(saved) // restores messages + config from loadSessionFile
agent.track(trackingId) // live snapshot of one sub-agent worker (`NAME-{TrackingID}`)
agent.listTrackedSubAgents() // TrackingIDs of known workers
agent.subscribeToSubAgent(id, fn) // live worker updates, returns unsubscribe fn
agent.persistNow() // rewrites the persist target now (no-op unless persist is set)prompt accepts string | ContentPart[].
Use ContentPart[] when sending images, audio, or video alongside text.
AgentRunOptions
| Field | Meaning / Use Case |
| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| stream | When true, returns a stream instead of a regular response promise. |
| onDelta(delta, event) | Called for each text chunk. Use to render answers live. |
| onThinkingDelta(delta, event) | Called for each reasoning chunk. Use to render reasoning separately. |
| onEvent(event) | Called for every event: text_delta, thinking_delta, tool_call_complete, tool_result, subagent_complete, subagent_delta, usage, done, and error. |
| wrapThinking | Wraps reasoning as <think>\n...\n</think>\n\n, allowing UIs to render it without maintaining separate state. |
| onTurn(turn) | Called after every model/tool turn with { turns, text?, thinking?, toolCalls?, toolResults? }. Drives sub-agent step logs and persistence. |
| additionalContext | Added only to the current user turn. Keeps the system prompt stable for better caching. |
| sessionId | Per-run session override. |
| headers | Per-run header merge. |
| signal | AbortSignal used to cancel generation, tools, and subagents. |
Streaming
const res = await agent.run("Design an append-only log", {
stream: true,
wrapThinking: true,
onThinkingDelta: (d) => process.stdout.write(d),
onDelta: (d) => process.stdout.write(d),
onEvent: (e) => {
if (e.type === "subagent_complete") {
console.log(`done: ${e.subagent!.name}`);
}
if (e.type === "tool_result") {
console.log(`tool: ${e.toolResult!.name}`);
}
},
});Manual Iteration
const stream = agent.stream("Hello");
for await (const e of stream) {
if (e.type === "text_delta") {
process.stdout.write(e.delta!);
}
if (e.type === "thinking_delta") {
process.stdout.write(e.thinkingDelta!);
}
}
const final = await stream.result();
stream.on("text_delta", (e) => console.log(e.delta));
stream.off("text_delta", handler);SubAgent
SubAgent extends Agent and is designed for fixed worker roles.
import { Agent, SubAgent } from "agent-accelerator";
const researcher = new SubAgent({
name: "researcher",
instructions: "You collect quantifiable metrics.",
model: "google/gemini-3.5-flash-lite",
});
const lead = new Agent({
model: "google/gemini-3.7-flash",
instructions: "Delegate research, then synthesize.",
subagents: [researcher],
});Differences from Agent
modelresolves asconfig.model ?? SUB_AGENT_MODEL ?? MODELand throws if empty.cachedefaults to{ retention: "short" }..asTool(name?, desc?)/.toTool(name?, desc?)converts the subagent into aToolDefinitionfor manual registration.stateless: trueis supported for one-shot judges that must not retain history.
SubAgentConfig is identical to AgentConfig.
Pre-defined vs Dynamic Delegation
Pre-defined delegation
Pass subagents: [a, b].
Each subagent is automatically registered as an internal tool on the parent that accepts:
{ task: string }Dynamic delegation
subagents lists pre-defined workers (each becomes a callable tool). For
LLM-spawned workers, configure dynamicSubagents:
const agent = new Agent({
name: "Main Agent",
model: process.env.MODEL,
tools: { get_topic_brief },
instructions: "You are the Main Agent",
dynamicSubagents: {
enabled: true,
maxSpawn: 4,
thinkingLevel: "low",
tools: { get_weather, recent_news },
timeout: 60000,
},
});A spawn_subagents tool is injected with the following task shape:
{
tasks: {
name,
role?,
instructions,
task,
tools?: string[],
timeoutMs?: number
}[]
}maxSpawn— max workers per call. Extras are trimmed safely, and the limit is stated in the tool description plus the Main Agent's system prompt so the model knows it.tools— developer-owned pool. Workers get zero tools unless the Main Agent grants a per-tasktoolssubset by name (unknown names are ignored), keeping worker context small.timeout(ms) —0= no limit,-1= the Main Agent sets a per-tasktimeoutMs,>0= fixed limit for every worker. Timed-out workers report an error entry; the rest of the batch still completes.- Workers are stateless: one task in, one result out, then shut down. No history, no recursion. The Main Agent cannot choose worker models or reasoning levels.
- Worker names use UPPER-KEBAB (
HBM-PRICING-ANALYST). Each worker gets a system-generated 32-char TrackingID, displayed asNAME-{TrackingID}(e.g.HBM-PRICING-ANALYST-9f2c…). - Workers stream internally: thinking/text deltas update the worker trace live (
agent.track(id),subscribeToSubAgent) and surface on the parent stream assubagent_deltaevents.
const res = await agent.run("Audit auth pipeline and write a threat model");
console.log(res.subagents.map((s) => s.name));Delegation Helpers
buildAgentTools(list)— wraps multiple agents and deduplicates names. Used internally bysubagents.createSubagentSpawnTool(parentAgent)— builds the dynamic subagent spawner. Rarely needed directly.DynamicSubagentTask—{ name, role?, instructions, task, tools?, timeoutMs? }, the shape produced for dynamic delegation. Nomodel— workers run on the developer-configured model.
Tools
import { Agent, tool, z } from "agent-accelerator";
const agent = new Agent({
model: "google/gemini-3.5-flash-lite",
tools: {
fetch_metrics: tool({
description: "Fetch cluster metrics",
input: z.object({
clusterId: z.string(),
}),
execute: async ({ clusterId }) => ({
clusterId,
load: 0.42,
}),
}),
},
});Tool Helpers
tool({ name?, description, input?, parameters?, strict?, timeoutMs?, maxTries?, maxConcurrency?, execute })— creates aToolDefinition. Provide either a Zodinputschema or raw JSONparameters.execute(input, ctx)may return arbitrary values; results are converted safely for model context.timeoutMsdefaults to0(no time limit) and is a best-effort event-loop deadline;ctx.signalenables cooperative cancellation.maxTriesaccepts a number or numeric string; a positive value is the total attempt limit, while0/omitted falls back to 3 attempts.maxConcurrencylimits simultaneous calls for that tool (default pool: 8).toStandardToolDeclarations(record | array)— converts tools to{ name, description, parameters }for providers.zodToJsonSchema(schema)— converts Zod schemas using native conversion with a fallback extractor.cleanJsonSchema(schema)— removes$schema,$defs, anddefinitions, resolves$ref, and preserves explicitadditionalProperties.executeToolCalls({ tools, toolCalls, agentName?, parallel?, signal?, sessionId? })— executes tool calls. Calls validate Zod input, use bounded parallelism, enforce per-tool timeouts, retry transient failures, recover namespace/camel-case tool-name aliases, and returnToolResultRecord[]. Missing tools and aborts are represented as error results rather than thrown.convert_document_to_markdown— built-in tool that converts document files (PDF/Word/PowerPoint/Excel/OpenDocument/RTF/EPUB/CSV, local path or URL) to Markdown for models without native document parsing. Requires the optional@firecrawl/anydocpeer.
The agent loop also blocks an identical tool name and argument set when the model requests it in the immediately following turn. The synthetic error is returned to the model so it can reuse the prior result or change its arguments.
Tool Timing and Deadlines
Every ToolResultRecord includes durationMs, measured in whole milliseconds with a minimum displayed value of 1ms. The value covers the tool execution attempt (including retry/backoff time), but excludes queue wait and input-schema validation.
timeoutMs is specified in milliseconds:
const get_status = tool({
name: "get_status",
description: "Check user authentication status.",
timeoutMs: 1,
input: z.object({ username: z.string() }),
execute: async ({ username }, ctx) => {
// Pass ctx.signal to cancellable APIs such as fetch.
return username === "Akshat Dwivedi" ? "Valid" : "Invalid";
},
});timeoutMs: 0 or an omitted value disables the deadline. Positive values are best-effort event-loop deadlines: JavaScript cannot forcibly stop arbitrary synchronous code, so asynchronous implementations should honor ctx.signal. A timeout is returned as an error tool result and is not retried indefinitely unless a positive maxTries is configured.
Tool timing is available after a run:
const response = await agent.run("Check status");
for (const result of response.toolResults) {
console.log(`${result.name}: ${result.durationMs}ms`);
}Tool Types
ToolExecutionContext:
{
toolCallId,
agentName?,
signal?,
sessionId?
}ToolCallRecord:
{
id,
name,
arguments,
rawArguments?,
thoughtSignature?,
callId?,
}ToolResultRecord:
{
id,
name,
result,
isError?,
durationMs?
}Providers
First-class provider IDs are:
googleopenrouteropenai
Any other provider prefix is treated as an OpenAI-compatible custom provider.
Model Strings
"google/<id>"→ Google AI Studio atgenerativelanguage.googleapis.com."openrouter/<scope>/<model>"or"scope/model:variant"→ OpenRouter."openai/<id>"→ OpenAI."groq/<id>","ollama/<id>", etc. → custom providers using{PREFIX}_API_KEYand{PREFIX}_BASE_URL. The API key is optional for local endpoints.- A bare
"some-model"usesOPENAI_BASE_URLwhen configured. Otherwise, catalog lookup is attempted, followed by anopenrouterfallback.
Model resolution never silently selects a default model.
An empty model throws.
An unknown prefix with no catalog match still routes to a custom provider, allowing private model IDs and endpoints.
Registry
resolveModel("google/x" | ModelSpec)→{ provider, modelId, modelSpec? }. Use to inspect routing before execution.getProvider("groq")→ returns a cached provider and automatically creates a custom provider for unknown prefixes.ensureCustomProvider(prefix, { baseUrl?, apiKey?, name? })→ gets or creates a custom provider. Useful for multiple endpoints in one process.normalizeProviderPrefix(s)→ lowercases and trims a provider prefix.ModelProvider.GoogleGenAI(model, apiKey?, { thinkingLevel?, baseUrl? })— same shape is available for.OpenRouter,.OpenAI,.Custom,.OpenAICompatible, and.Generic.
Each returns a ModelProviderInstance:
{
model,
apiKey?,
baseUrl?,
thinkingLevel?
}The result can be passed directly to Agent.model.
Use Custom for Groq, Together, Ollama, vLLM, and similar endpoints.
import { Agent, ModelProvider } from "agent-accelerator";
new Agent({
model: ModelProvider.GoogleGenAI("gemini-3.5-flash-lite"),
});
new Agent({
model: ModelProvider.Custom(
"groq/llama-3.3-70b-versatile",
process.env.GROQ_API_KEY,
{
baseUrl: process.env.GROQ_BASE_URL,
},
),
});Provider is the provider contract containing id, name, models, getModel, generate, and stream.
Its catalog-first getModel behavior includes a permissive fallback (createGenericModelSpec) so private model IDs never fail preflight.
Implement the Provider interface when adding a fully custom transport.
ProviderRequestOptions { apiKey?, baseUrl?, headers?, thinking?, cache?, serviceTier?, tools?, toolChoice?, signal?, sessionId?, env? } is the per-call options object passed to generate/stream. Agent builds it automatically.
GoogleInteractionsProvider (aliased as GoogleAIStudioProvider) and GOOGLE_MODELS provide Google support over the Interactions API (POST {baseUrl}/interactions, streaming via ?alt=sse).
Turns chain statefully through previous_interaction_id per session, falling back to stateless full-history sends when the session switches providers or models.
Thinking levels map to thinking_level, with thinking_summaries enabled whenever thinking is active:
minimal | low | medium | high (xhigh/max clamp to high; none is unsupported and omitted; dynamic omits)Google-specific handling includes:
- Isolating
thoughtSignatureper part for prefix stability. - Converting JSON Schema to OpenAPI 3.0 through
stripSchemaForGoogle. - Implicit/automatic caching only — explicit retention and
cachedContentIdare warned about and dropped (the Interactions API defines no explicit cache primitives).
Helpers:
createExplicitCache({ model, systemInstruction?, contents?, tools?, displayName?, ttlSeconds?, expireTime?, apiKey?, baseUrl? })isValidThoughtSignature(sig)retainThoughtSignature(existing, incoming)extractGoogleThoughtSignature(obj)stripSchemaForGoogle(schema)clearInteractionChains(sessionId?)
OpenAI
OpenAIResponsesProvider (aliased as OpenAIProvider) and OPENAI_MODELS provide OpenAI support over the Responses API (POST {baseUrl}/responses).
Every turn is stateless: the full history is sent explicitly with store: false, so no previous_response_id, background, or conversation chaining is ever used.
Thinking levels map to reasoning.effort verbatim:
none | minimal | low | medium | high | xhigh | max (dynamic omits — server default)It also:
- passes
service_tier(flex|priority); - pins cache affinity with
prompt_cache_key(clamped to 64 chars); - rejects
videoparts up front with a one-line error (the Responses wire carries text/image/file only); - never sends
temperature, token caps, or background/conversation fields.
OpenRouter
OpenRouterChatCompletionsProvider (aliased as OpenRouterProvider / OpenRouter) and OPENROUTER_MODELS provide OpenRouter support over Chat Completions (POST {baseUrl}/chat/completions).
Every turn is stateless: the full messages[] array is sent explicitly — no chaining primitives exist on this endpoint.
Session affinity is a top-level body session_id plus the x-session-id header fallback (body takes precedence per OpenRouter docs).
Thinking levels map to reasoning.effort verbatim:
none | minimal | low | medium | high | xhigh | max (dynamic omits — server default)It also:
- passes
service_tier(flex|priority); - enables the
file-parserplugin when file parts are present (remote file URLs are additionally surfaced in text); - maps
tool_choice(autoomitted;none,required, and function pins supported); - reports
cached_tokens/cache_write_tokens/ reasoning tokens, preferring provider-reportedcostwhen present.
serviceTier values other than flex / priority are omitted (standard routing).
Custom
OpenAICompatibleProvider is also exposed through the CustomProvider, createCustomProvider, and createOpenAICompatibleProvider(prefix, opts?) aliases.
CustomProviderOptions:
{
name?,
baseUrl?,
apiKey?,
defaultBaseUrl?
}createGenericModelSpec(provider, modelId) creates a permissive placeholder so private model IDs can pass validation and window checks.
Provider wire types are exported for inspecting raw requests and responses through:
res.raw.request.body
res.raw.response.bodyExamples include:
OpenAIChatCompletionRequestGoogleGenerateContentRequestOpenRouterChatRequest
Thinking
Thinking uses a single thinkingLevel flag:
ThinkingLevel =
| "none"
| "dynamic"
| "minimal"
| "low"
| "medium"
| "high"
| "xhigh"
| "max";Thinking Levels
none— disables thinking where supported.dynamic— allows the model to decide.minimal,low,medium,high,xhigh,max— request increasing levels of reasoning where supported.
Thinking can be configured on the Agent or through ModelProvider.*(..., { thinkingLevel }).
The per-run thinkingLevel option can override the agent default for one request.
Internally, thinking is normalized to:
ThinkingConfig {
enabled?,
level?,
budgetTokens?,
includeThoughts?
}Catalog Helpers
getModelThinkingInfo(provider, modelId)→{ supportsThinking, reasoningOptions?, allowedLevels, supportsDisable, description }. Useful for building UI selectors.validateModelThinking(provider, modelId, level)→ throwsThinkingLevelError { provider, modelId, requestedLevel, allowedLevels, supportsThinking }with a fix hint. Automatically called byrunandstream. Catch it and useerr.allowedLevelsto offer valid options.resolveEffectiveThinking(base, override?)→ resolves a per-runthinkingLeveloverride into aThinkingConfigwithout mutating agent config.getModelsForProvider("google")→ returns known model specifications.
Unknown models allow all thinking levels.
Fixed-reasoning models reject none.
Cache
CacheConfig {
retention?,
cachedContentId?,
sessionId?,
ttlSeconds?
}CacheRetention =
| "implicit"
| "short"
| "medium"
| "long";Retention
implicit— automatic prefix reuse without an explicit cache object or storage fee. This is the only retention the native adapters act on (by doing nothing special — stable prompts plus session affinity do the work).short/medium/long— TTL hints (5m/1h/12h) consumed only by the opt-increateExplicitCachehelper. Provider adapters warn about and drop any non-implicitretention: none of the native transports expose retention control.
Session Affinity
sessionId pins provider affinity using mechanisms such as:
x-session-id(+x-client-request-id) headers — all providers, best effort- top-level body
session_id— OpenRouter (takes precedence over the header) prompt_cache_key— OpenAI only (custom endpoints get headers only; strict ones reject unknown body fields)
Reuse the same session ID across turns when cache affinity is desired.
In browser runtimes custom x-* headers are stripped to avoid CORS preflights: OpenAI keeps affinity through its body key, while custom endpoints lose affinity entirely on browsers.
Explicit Caches
createExplicitCache mints a Google cachedContents/... resource directly (TTL via retention/ttlSeconds).
cachedContentId is accepted and carried in context, but no native adapter currently sends it — explicit references are warned about and dropped, so stable prompts plus session affinity remain the cache-reuse path.
ttlSeconds overrides the retention mapping.
Provider Cache Mechanisms
- Google keeps system prompts and tools stable and isolates signatures (implicit caching only).
- OpenAI pins affinity with
prompt_cache_key. - OpenRouter pins affinity with body
session_idplus headers (sticky routing from the first request). - Custom endpoints get headers-only affinity; anything else cache-related is warned about and dropped.
The following cache helpers are exported:
getPromptCacheRetentionretentionToTtlSecondsclampCacheKey
For the best cache reuse, keep instructions and the tool set stable. Put changing data in user messages or additionalContext.
serviceTier
serviceTier = "flex" | "priority";Omit serviceTier for standard routing.
Where supported, it is passed as service_tier or translated into provider-specific routing.
flex— suitable for batch-tolerant workloads.priority— suitable for latency-sensitive workloads.
Streaming and Events
AssistantMessageEventStream implements AsyncIterable<StreamEvent>.
It provides:
.result(): Promise<AgentResponse>
.push(...)
.end(...)
.fail(...)
.on(...)
.off(...)
.cancel()
.onCancel(...)
.isCancelled()StreamEvent
StreamEvent {
type,
delta?,
thinkingDelta?,
toolCall?,
toolResult?,
subagent?,
subagentTrackingId?,
usage?,
responseId?,
finishReason?,
error?,
partialText?,
partialThinking?,
raw?
}Event types:
start
text_start
text_delta
text_end
thinking_start
thinking_delta
thinking_end
tool_call_start
tool_call_delta
tool_call_complete
tool_result
subagent_complete
subagent_delta
usage
done
errorEmitted in practice: start, text_delta, thinking_delta, tool_call_complete, tool_result, subagent_complete, subagent_delta, usage, done, error. The rest exist in the type for forward compatibility.
Use:
text_deltafor answer output.thinking_deltafor reasoning output.tool_call_complete/tool_resultfor tool progress.subagent_completefor each completed worker during streaming multi-agent runs.subagent_deltafor live worker thinking/text (subagentTrackingId+delta/thinkingDelta+ partials). Render per-worker; never merge into the parent answer.usagefor interim usage counts.donefor final usage and completion information.
Breaking out of for await auto-cancels. Cancellation joins your
AbortSignal and stream cancellation into one linked controller, so
controller.abort() and stream.cancel() both stop the HTTP request.
Abort rejects with AbortError (never retried) and never resolves partial
results. agent.run(prompt, { stream: true, signal }) exposes the same
cancel handle.
SSEParser provides:
feed(chunk)
flush()It parses provider SSE streams and is only needed when implementing custom transports.
Response, Usage, and Raw Data
AgentResponse
AgentResponse {
text,
thinking?,
thoughtSignature?,
toolCalls,
toolResults,
subagents,
usage,
responseId?,
model,
provider,
finishReason?,
durationMs,
raw,
turns
}AgentResponse also provides:
.toString()
.toJSON()SubAgentExecutionMetadata
SubAgentExecutionMetadata {
name,
role?,
task,
model,
provider,
durationMs,
usage,
turns,
finishReason?,
responseId?,
text,
thinking?,
toolCalls?,
raw?,
isError?,
error?,
trackingId?,
sessionId?,
parentSessionId?,
steps?
}finishReason is "max_turns" when a finite maxTurns limit stopped the run with tool calls still pending (plus a console warning naming the dropped tools).
SubAgentStep
SubAgentStep {
turn,
type, // "assistant" | "tool_call" | "tool_result"
name?,
text?,
thinking?,
isError?,
timestamp,
partial? // true while the worker is still streaming this step
}agent.track(id) returns these live; partial entries finalize when the worker's turn completes.
TokenUsage
TokenUsage {
inputTokens,
outputTokens,
totalTokens,
cachedTokens?,
cacheReadTokens?,
cacheWriteTokens?,
thinkingTokens?,
cost?: {
inputCost?,
outputCost?,
cacheReadCost?,
cacheWriteCost?,
totalCost?
}
}Cost calculation prefers provider-reported totals. When unavailable, costs are computed using catalog pricing.
Subagent usage is rolled into the parent totals.
ProviderRawData
ProviderRawData {
request: {
url,
method,
headers,
body
},
response?: {
status,
statusText,
headers,
body
}
}Sensitive keys are redacted.
Errors
Provider failures throw AgentAccelProviderError — one actionable line,
never a wire dump:
[openrouter/nvidia/nemotron-3.5-lightning:free] request failed (404): No endpoints found that support input videoRaw details (statusCode, url, truncated body) stay attached as
non-enumerable properties: available programmatically, invisible in runtime
dumps. URL query strings are stripped (keys sometimes live there); headers
and cookies are never attached. Helpers:
toConciseProviderError(err, providerId, modelId)— collapse any provider error.assertModalitiesSupported(context, providerId, modelId)— pre-request modality gate.
Examples print failures as exactly one line via examples/_shared.ts:
const res = await agent.run([...]).catch(fail);
// ✖ [provider/model] request failed (404): ...
// exit 1, no stack dumpModel Catalog & Dynamic 12-Hour TTL Sync
Agent Accelerator features an adaptive, dynamically synchronized model catalog powered by models.dev. To avoid shipping a bloated 4.4 MB static JSON file with production builds, the SDK employs a high-performance 12-hour TTL local caching architecture:
- Automated 12-Hour Cache Validation: When
.run(),.ask(), or.stream()executes, the runtime checkssrc/data/models.dev.json. If the cache timestamp is within 12 hours, it reads from disk with zero network delay. When the TTL expires, it transparently synchronizes withhttps://models.dev/api.json. - Git & Package Safety: The dynamic cache file (
src/data/models.dev.json) is gitignored and excluded from production packages.
Developer Catalog Controls
import {
refreshModelCatalog,
getCatalogStatus,
setCatalogTTL,
getModelFromCatalog,
getModelThinkingInfo,
} from "agent-accelerator";
// Force an immediate refresh from models.dev:
await refreshModelCatalog({ force: true });
// Customize default TTL (e.g. 24 hours):
setCatalogTTL(24 * 60 * 60 * 1000);
// Inspect cache health and metadata:
const status = getCatalogStatus();
console.log(status.modelCount, status.providerCount, status.isExpired);CLI Refresh
bun run update-models # Refresh model catalog (force or expired)
bun src/update-models.ts --force # Force re-download
bun src/update-models.ts --ttl=24h # Refresh with custom TTLUpstream data lags on some inputs (e.g. gpt-4o accepts audio). Record
verified corrections in MODALITY_OVERRIDES in src/models/catalog.ts
(keyed provider/model, merged over the snapshot) — never edit the JSON
directly, a refresh would wipe it.
Catalog Helpers
getModelFromCatalog(provider, modelId)→ModelSpec | undefined, including alias and global fallback handling. Use it for context-window, pricing, and modality checks.getModelsForProvider(provider)→ returnsModelSpec[].getModelThinkingInfo(provider, modelId)→ inspect permitted thinking levels.refreshModelCatalog(options?)→ programmatically sync with upstream models.dev.getCatalogStatus()→ diagnostic info on active catalog, provider count, and TTL expiry.
ModelSpec
ModelSpec {
id,
provider,
name,
description?,
family?,
api?,
contextWindow,
maxOutputTokens,
limit?,
cost?,
modalities?,
reasoning?,
reasoning_options?,
tool_call?,
capabilities,
pricing?,
raw?
}contextWindow and maxOutputTokens mirror:
limit.context
limit.outputModelCapabilities
ModelCapabilities {
supportsThinking?,
supportsThinkingBudget?,
supportsThinkingLevel?,
supportsImplicitCaching?,
supportsExplicitCaching?,
supportsLongCacheRetention?,
supportsParallelToolCalls?,
supportsStreaming?,
modalities?,
supportsReasoningToggle?,
supportsReasoningEffort?
}Messages and Media
Message
Message {
role: "system" | "user" | "assistant" | "tool",
content: string | ContentPart[],
name?,
thoughtSignature?
}ProviderContext
ProviderContext {
systemPrompt?,
messages,
tools?,
cachedContentId?
}ContentPart
Supported content parts include:
TextPart { type:"text", text, thoughtSignature? }— plain text.ThinkingPart { type:"thinking", thinking, thoughtSignature? }— preserved reasoning.ToolCallPart { type:"tool_call", id, name, arguments, rawArguments?, thoughtSignature? }ToolResultPart { type:"tool_result", id, name, result, isError? }ImagePart { type:"image", image, mimeType? }AudioPart { type:"audio", audio, mimeType? }VideoPart { type:"video", video, mimeType? }FilePart { type:"file", file, mimeType?, filename? }— PDFs/documents.
image, audio, video, and file inputs accept:
- data URLs
- remote
http(s)URLs - local paths
- raw base64
Uint8ArrayArrayBuffer
normalizeMediaInput(input, mime?) returns:
{
mimeType,
base64Data,
dataUrl
}inferMimeType(path) infers the MIME type from a file extension.
Media normalization is handled automatically by providers (mapped to provider-native media shapes). Thinking traces are echoed in follow-up turns only where accepted — strict endpoints receive tool calls without reasoning_content.
Provider matrix: text + image + wav/mp3 audio + PDF work on all supported providers. Video works on Gemini only.
Before any network call, the executor checks image / audio / video /
file-as-PDF parts against the catalog's modalities.input for that
provider/model and fails fast with a one-line error naming the gap
(assertModalitiesSupported). Models absent from the catalog (custom
providers, dynamic routers) are skipped — the provider endpoint decides and
its verdict surfaces as a concise error (see Errors). Verified
catalog corrections live in MODALITY_OVERRIDES (src/models/catalog.ts),
never in the gitignored snapshot.
Utils
createSessionId(prefix="accel")— creates a UUID-based session ID clamped to 64 characters for affinity.createTrackingId()— system-generated 32-char TrackingID for one sub-agent worker (NAME-{TrackingID}).createChildSessionId(parentId, tag)/createFixedChildSessionId(parentId, name)/isSessionDescendant(child, parent)— provider-safe (≤64 chars) child session IDs with parent-hash lineage.getEnv(key, fallback?)— resolves an environment variable.getApiKey(provider, explicit?, env?)— resolves API keys using the configured environment lookup order.buildSessionHeaders(provider, cache?, custom?, sessionId?)— builds provider-specific affinity headers. Normally handled automatically.withRetries(fn, { maxRetries?, maxRetryDelayMs?, signal?, label? })/isTransientError(err)— bounded retries for transient provider failures (defaults: 2 retries, 5s cap, aborts never retried).convertDocumentToMarkdown(input, { filename?, format?, mimeType?, maxChars? })— converts PDF/Word/PowerPoint/Excel/OpenDocument/RTF/EPUB/CSV to Markdown via the optional@firecrawl/anydocpeer. ThrowsDocumentConversionErroron failure.convert_document_to_markdown— built-in model-callable version of the above (returns a<Document>envelope); auto-registered withbypassInputFileModality: true.z— re-exported from Zod so tools do not require a separate Zod import.
Examples
bun run examples/05-chat.ts
# persistent CLI, /model "..." /level <lvl> /help /exit
bun run examples/04-sub-agents.ts
# fixed researcher + critic pipeline
bun run examples/03-multi_agent.ts
# dynamic spawn_subagents demo
bun run examples/01-metadata.ts
# usage + raw inspection
bun run examples/02-function_calling.ts
# single tool call
bun run examples/06-multimodal_image.ts [./photo.png]
# image input (remote URL default, local path optional)
bun run examples/07-multimodal_audio.ts
# audio input
bun run examples/09-multimodal_document.ts
# PDF/file input
bun run examples/08-multimodal_video.ts
# video input (video-capable model required)
bun run examples/10-document_markdown.ts
# document → Markdown preprocessing (any model, optional @firecrawl/anydoc peer)
bun run examples/11-session-identity.ts
# sub-agent TrackingIDs, live worker boxes, lineage, track(), durable sessionsEvery example ends its run() with .catch(fail) (examples/_shared.ts),
so failures print one line and exit 1 — no stack dumps.
The chat example persists conversations to:
sessions/<sessionId>/session.json (+ media/ for images, audio, video, files)It resumes the most recent session on next launch (legacy .session.json/.session.jsonl files still load) and prints a resume banner:
↺ Previous session loaded • accel-1a2b… (.session.json) • 12 messages • google/gemini-3.5-flash-lite • total-CH82.4% • $0.013 totalIt also displays per-turn and session totals on one line:
↑turn ↓turn CRturn turn-CH% | ↑total ↓total CRtotal total-CH% $total(+turn) ctx%/windowSession persistence
import {
Agent,
loadSessionFile,
saveSessionFile,
SessionTelemetry,
formatSessionBanner,
} from "agent-accelerator";
const saved = loadSessionFile(".session.json");
const telemetry = SessionTelemetry.fromSaved(saved);
const agent = new Agent({
model: saved?.model ?? process.env.MODEL,
sessionId: saved?.sessionId,
});
if (saved) agent.importSession(saved);
const res = await agent.run("Hello");
telemetry.add(res.usage, agent.modelStringOrSpec as string);
saveSessionFile(".session.json", agent, telemetry);agent.exportSession(telemetry)/agent.importSession(saved)round-trip messages, system prompt, model, thinking level, instructions, cache, worker model, and sub-agent traces without touchingAgentContextinternals.SessionTelemetryaccumulatesinput / output / cacheRead / cacheWrite / reasoning / cost, with clampedturnHitRate()/totalHitRate(), a dualformatBar(res), andformatSessionBanner(saved, telemetry, path)for startup.- Files are pretty-printed JSON (2-space indent), written atomically (temp + rename). Cost prefers provider-reported totals and falls back to catalog pricing.
- Binary media (bytes, data URLs, base64) is extracted to
media/on save viasaveSessionDir("sessions", agent, telemetry)→sessions/<sessionId>/{session.json, media/*}; remote URLs and local paths stay references.loadSessionDir(dir)resolvesmedia/…refs back to absolute paths, andfindLatestSessionDir("sessions")resumes the most recent session. - Sub-agent traces persist in a dedicated
subagentssection keyed by TrackingID ({ trackingId, name, sessionId, parentSessionId, status, task, turns, usage, steps, text }). Files stay lean: redundantrawArgumentsand repeated step text are dropped on save (rebuilt/kept live in memory). - With
persist: { dir }or{ file }, the session file is rewritten after every step, so a crash loses at most the in-flight step.
Scripts and Structure
Scripts
bun run typecheck # tsc --noEmit
bun test # bun test test/
bun run update-models # refresh model catalog cache (supports --force, --ttl=24h)Project Structure
src/
├── agent/ # Agent, context, loop, delegation, subagent
├── session/ # Session persistence + telemetry (snapshots, hit rates, .session.json store)
├── providers/ # native REST adapters (google/openai/openrouter/openai-compat) + registry + canonical contract
├── models/ # Dynamic catalog cache, parser, verified overrides
├── data/ # Dynamic model catalog cache (gitignored, excluded from bundle)
├── tools/ # tool(), schema, executor
├── streaming/ # event stream, SSE parser
├── types/ # agent, core, message, model, response, tool
└── utils/ # base64, cache, env, headers, media, serialization, session, thought-signature, documents, retry, errors
examples/
├── 01-metadata.ts 02-function_calling.ts 03-multi_agent.ts
├── 04-sub-agents.ts 05-chat.ts 06-multimodal_image.ts
├── 07-multimodal_audio.ts 08-multimodal_video.ts
├── 09-multimodal_document.ts 10-document_markdown.ts
├── files/ prompts/ research-agent.ts _shared.ts
test/ # unit + mocked-provider tests mirroring src/License
MIT License — see LICENSE for details.
Copyright (c) 2026 Sashvat Bharat.
