@stackfactor/agent-utils
v1.3.1
Published
Shared utilities for StackFactor AI agent services — LangChain helpers, structured logging, error handling, and constants.
Keywords
Readme
@stackfactor/agent-utils
Shared utilities for StackFactor AI agent services — LangChain helpers, structured logging, error handling, and constants.
Installation
npm install @stackfactor/agent-utilsModules
The package exports four modules:
import {
langChain,
logger,
errorHandling,
constants,
} from "@stackfactor/agent-utils";langChain
Unified interface for running LLM prompts, managing LangChain agents, and generating images across OpenAI, Anthropic, Google, DeepSeek, Kimi (Moonshot), and GLM (Zhipu) providers.
langChain.runPromptWithModel(modelName, config, prompt, onProgressReport?, minPercent?, maxPercent?, expectsJsonResponse?, schema?, agentName?, tools?, usageTracker?)
Sends a prompt to an LLM and returns the response. Supports three execution modes:
- Agentic (
config.agentic === true) — creates and runs a LangChain agent with tools. - Streaming (
onProgressReportprovided) — streams the response with progress callbacks. Uses nativeresponse_formatfor OpenAI+schema, or NDJSON for other providers. - Non-streaming — direct invocation with JSON extraction.
When expectsJsonResponse is true (the default), JSON escape instructions are injected and the response is parsed. If a Zod schema is provided, the parsed result is validated.
| Parameter | Type | Default | Description |
| --------------------- | -------------------- | --------------- | -------------------------------------------------------------------------------------------------------- |
| modelName | string | — | Model identifier: "gpt-4o", "claude-3-5-sonnet", "gemini-1.5-pro", "deepseek-chat", "kimi-k2-0905-preview", "glm-4.6", etc. |
| config | object | — | API keys (openAIAPIKey, anthropicAPIKey, googleAPIKey, deepSeekAPIKey, kimiAPIKey, glmAPIKey), temperature, agentic, recursionLimit, tavily, tavilyAPIKey |
| prompt | string \| object[] | — | Plain string or array of { role, content } message objects |
| onProgressReport | function \| null | null | Async callback receiving { message, progress } updates |
| minPercent | number | 0 | Lower bound for progress percentage |
| maxPercent | number | 100 | Upper bound for progress percentage |
| expectsJsonResponse | boolean | true | Parse response as JSON |
| schema | ZodSchema \| null | null | Zod schema for response validation |
| agentName | string | "StackFactor" | Display name for the agent (agentic mode) |
| tools | any[] | [] | LangChain tools available to the agent (agentic mode) |
| usageTracker | UsageTracker \| null | null | Optional accumulator updated on every successful LLM call with cost (USD) and per-model token counts. |
Returns: JSON string (when expectsJsonResponse is true) or raw content string.
const usageTracker = {};
const result = await langChain.runPromptWithModel(
"gpt-4o",
{ openAIAPIKey: "sk-..." },
"Generate a summary of this document.",
null,
0,
100,
true,
null,
"StackFactor",
[],
usageTracker,
);
// usageTracker is now:
// {
// cost: 0.00123,
// tokens: {
// "gpt-4o_inputTokens": 412,
// "gpt-4o_outputTokens": 87,
// }
// }langChain.createAIAgent(name, modelName, systemPrompt, tools?, responseFormat?, config, onReportProgress?, minPercent?, maxPercent?)
Constructs a LangChain agent with a model, system prompt, and tools. When onReportProgress is provided, a report_progress tool is automatically added.
| Parameter | Type | Default | Description |
| ------------------ | ------------------ | ------- | ---------------------------------------- |
| name | string | — | Display name for the agent |
| modelName | string | — | LLM identifier |
| systemPrompt | string | — | System prompt describing agent behaviour |
| tools | any[] | [] | LangChain tool instances |
| responseFormat | any | — | Structured response format descriptor |
| config | object | — | API keys, temperature, tavily, etc. |
| usageTracker | object \| null | null | Accumulator the Tavily tools bill credits and record sources into |
| onReportProgress | Function \| null | null | Progress callback |
| minPercent | number | 0 | Minimum reportable progress |
| maxPercent | number | 100 | Maximum reportable progress |
Returns: A configured LangChain agent instance.
langChain.runAIAgent(agent, prompt, config, onProgress?)
Executes a LangChain agent with a user prompt. Registers callbacks for tool_start, tool_end, agent_action, and error events when onProgress is provided. Logs execution time on completion.
| Parameter | Type | Default | Description |
| ------------ | ------------------ | ------- | -------------------------------- |
| agent | any | — | Agent created by createAIAgent |
| prompt | string | — | User message to send |
| config | object | — | recursionLimit (default: 25) |
| onProgress | Function \| null | null | Progress event callback |
Returns: Raw response from the agent's invoke method.
const agent = langChain.createAIAgent(
"Summarizer",
"gpt-4o",
"You summarize documents concisely.",
[],
null,
{ openAIAPIKey: "sk-..." },
);
const response = await langChain.runAIAgent(
agent,
"Summarize this text...",
{},
);langChain.runPromptWithModelForImageGeneration(modelName, config, prompt, options?)
Generates an image using OpenAI or Google AI models.
| Parameter | Type | Default | Description |
| ----------- | -------- | ------- | --------------------------------------------------------------------------------------------------------- |
| modelName | string | — | Image model: "dall-e-3", "gpt-image-1.5", "imagen-4.0-generate-001", "gemini-3.0-pro-image", etc. |
| config | object | — | openAIAPIKey or googleAPIKey |
| prompt | string | — | Text prompt describing the image |
| options | object | {} | Provider-specific options (see below) |
OpenAI options: size ("1024x1024", "1792x1024", etc.), style ("vivid" or "natural"), responseFormat ("url" or "b64_json"), n (number of images).
Google options: aspectRatio ("1:1", "3:4", "4:3", "9:16", "16:9"), numberOfImages, negativePrompt.
Returns: { url?, b64_json?, revisedPrompt? } for single images, or { images: [...] } for multiple.
const image = await langChain.runPromptWithModelForImageGeneration(
"dall-e-3",
{ openAIAPIKey: "sk-..." },
"A futuristic city skyline at sunset",
{ size: "1792x1024" },
);langChain.throwErrorIfNotSuccessful(response)
Guards that a response is a non-empty string. Throws an INTERNAL_SERVER_ERROR if not.
| Parameter | Type | Description |
| ---------- | ----- | ----------------- |
| response | any | Value to validate |
Returns: The response string if valid.
logger
GCP-compatible structured logging via Winston with OpenTelemetry trace enrichment.
logger.log(request, level, message, options?)
Writes a structured log entry. Automatically enriches with OpenTelemetry traceId and spanId when an active span exists. Prepends the user's email from the request object when available.
| Parameter | Type | Default | Description |
| --------- | ---------- | ------- | -------------------------------------------------------------------------- |
| request | any | — | HTTP request object (reads request.user.email), or null |
| level | LogLevel | — | "error", "warn", "info", "http", "verbose", "debug", "silly" |
| message | string | — | Log message |
| options | object | {} | Additional structured fields (service, requestId, etc.) |
logger.log(req, logger.levels.info, "User signed in", { service: "auth" });
logger.log(null, logger.levels.error, "Connection failed");logger.levels
Enum-like object mapping level names to their string values:
logger.levels.error; // "error"
logger.levels.warn; // "warn"
logger.levels.info; // "info"
logger.levels.http; // "http"
logger.levels.verbose; // "verbose"
logger.levels.debug; // "debug"
logger.levels.silly; // "silly"errorHandling
Factory for consistently shaped error payloads.
errorHandling.create(errorCode, errorMessage, details?)
Creates a structured error object with a stack trace.
| Parameter | Type | Default | Description |
| -------------- | -------- | ------- | ------------------------------------------ |
| errorCode | number | — | HTTP status code or application error code |
| errorMessage | string | — | Human-readable error description |
| details | any | null | Optional supplementary information |
Returns: { code, message, details, stack }
throw errorHandling.create(404, "User not found");
throw errorHandling.create(400, "Validation failed", { field: "email" });constants
Shared constants used across the package.
constants.HTTP_CODES
| Constant | Value |
| ----------------------- | ----- |
| BAD_REQUEST | 400 |
| UNPROCESSABLE_ENTITY | 422 |
| INTERNAL_SERVER_ERROR | 500 |
| BAD_GATEWAY | 502 |
constants.ERROR
| Constant | Value |
| ---------------------------- | ---------------------------------------- |
| UNABLE_TO_GENERATE_CONTENT | "Unable to generate content" |
| UNEXPECTED_ERROR | "An unexpected error occured..." |
| UNSUPPORTED_MODEL | "The specified model is not supported" |
Web Search (Tavily)
Set config.tavily to give an agent the Tavily web tools. These run client-side: the model emits a tool call, this library performs the HTTP request, and the result is appended to the conversation. Only an agent loop can execute them, so config.agentic must be true — enabling tavily without it throws BAD_REQUEST rather than handing the model tools whose calls nothing answers.
| Tool | Arguments | Returns | Default |
| ------------- | -------------------------------- | ----------------------------------------------------------------------- | ------- |
| web_search | query, maxResults? | Snippets and URLs; no page content | on |
| web_extract | urls[], query? | Page content — matching passages with a query, whole document without | on |
| web_map | url, instructions? | URLs only, no content | off |
| web_crawl | url, instructions?, limit? | Full content of every page visited | off |
No new dependency is added: all four endpoints are a single POST, issued with the runtime's own fetch.
Both reading modes, one tool
web_extract covers whole-page and targeted reading through Tavily's own optional query field, so the model chooses per call:
queryomitted → the entire document.querysupplied → only the matching chunks (chunks_per_source, default3).
Targeted extraction is dramatically cheaper in context and is usually sufficient; whole-document extraction is there for when the full structure matters. Splitting these into two tools would duplicate the description tokens on every request to express a distinction the API already makes with one optional field.
const usageTracker = {};
const result = await langChain.runPromptWithModel(
"claude-opus-5",
{
anthropicAPIKey: "sk-ant-...",
tavilyAPIKey: "tvly-...",
agentic: true, // required
"tavily-credit-costs": 0.008, // USD per credit
tavily: {
tools: ["search", "extract", "map"],
includeDomains: ["sc.gov", "sc.edu"],
maxCharsPerUrl: 60000,
},
},
"What are the SC procurement thresholds for sole-source awards?",
null, 0, 100, false, null, "StackFactor", [], usageTracker,
);
// usageTracker.tokens.tavilyCredits === 4
// usageTracker.cost includes 4 * 0.008
// usageTracker.webSearchSources === [{ url: "https://...", title: "..." }, ...]TavilyConfig
tavily: true uses the defaults below. Pass an object to configure it.
| Field | Default | Description |
| ----------------- | ------------------------ | -------------------------------------------------------------------------------------------- |
| tools | ["search", "extract"] | Which tools to expose; each one spends its description tokens on every request |
| searchDepth | Tavily's "basic" | "basic" \| "advanced" \| "fast" \| "ultra-fast" |
| extractDepth | Tavily's "basic" | "advanced" also pulls tables and embedded content, at double the credits |
| maxResults | 5 | Results per search |
| chunksPerSource | 3 | Relevant chunks per URL when a query is passed to web_extract |
| maxCharsPerUrl | 60000 | Per-URL cap before truncation (≈15k tokens) |
| maxCharsTotal | 150000 | Cap across one tool call (≈37k tokens) |
| includeDomains | — | Restrict web_search to these domains |
| excludeDomains | — | Exclude these domains from web_search |
| crawlLimit | 10 | Maximum pages one web_crawl or web_map call may return |
| format | "markdown" | "markdown" preserves table and heading structure at fewer tokens than the prose equivalent |
Context management
Nothing sits between Tavily's response and the context window — there is no provider-side filtering here — so the caps above are the only thing standing between one extract call and a 100-page PDF. Three measures apply by default:
Responses are stripped before the model sees them. Tavily returns score, published_date, favicon and id alongside each result. None of it informs an answer, so only the title, URL and text are forwarded; the rest is dropped rather than billed as input tokens.
web_search never requests page content. include_raw_content is always false: search finds documents, web_extract reads them. Asking for full text at search time pays for every result to find one.
Truncation is reported, never silent. When a document exceeds maxCharsPerUrl, or a call exhausts maxCharsTotal, the payload says exactly what was cut:
[truncated: 60,000 of 412,336 characters shown — narrow the query or request fewer URLs to see the parts that matter]A model that cannot tell it received half a regulation will reason over the half it got, which on a compliance corpus is worse than an error.
web_crawl is off by default for the same reason: it returns full content for every page it visits, and is the easiest way here to exhaust a context window. Prefer web_map to enumerate a site, then web_extract with a query on the few URLs that matter.
Notes
- Works with every supported provider, including DeepSeek, Kimi and GLM, which have no native web search of their own.
- Credits come from each response's own
usageblock rather than an estimate, and are accumulated onusageTracker.tokens.tavilyCredits, costed from thetavily-credit-costsconstant (USD per credit; falls back to Tavily's list price of0.008). - URLs are collected on
usageTracker.webSearchSources, deduplicated across the whole run — Tavily's terms require citing the original sources when its content is shown to end users. - Per-URL failures are surfaced in the tool result instead of being dropped, so a fetch failure is not mistaken for an absence of content.
- Tool calls honour the caller's abort signal, so a cancelled run stops paying for in-flight crawls.
Supported Models
Text / Chat
| Provider | Model Prefix | Example |
| ----------------- | ---------------------- | ------------------------------------ |
| OpenAI | gpt- | gpt-4o, gpt-4o-mini |
| Anthropic | claude- | claude-3-5-sonnet, claude-3-opus |
| Google | gemini- | gemini-1.5-pro, gemini-2.0-flash |
| DeepSeek | deepseek- | deepseek-chat, deepseek-reasoner |
| Kimi (Moonshot) | kimi-, moonshot- | kimi-k2-0905-preview, moonshot-v1-8k |
| GLM (Zhipu) | glm- | glm-4.6, glm-4.5 |
DeepSeek, Kimi, and GLM are routed through their OpenAI-compatible endpoints. Native structured output via response_format with JSON schema is only applied to gpt-* models; all other providers (including these three) use prompt-based JSON instructions.
Image Generation
| Provider | Model Prefix | Example |
| -------- | ----------------------- | ------------------------------------------------- |
| OpenAI | dall-e-, gpt-image- | dall-e-3, gpt-image-1.5 |
| Google | imagen-, gemini- | imagen-4.0-generate-001, gemini-3.0-pro-image |
Config Object
The config object accepted by LangChain methods supports the following keys:
| Key | Type | Description |
| ----------------- | --------- | ------------------------------------------- |
| openAIAPIKey | string | OpenAI API key |
| anthropicAPIKey | string | Anthropic API key |
| googleAPIKey | string | Google AI API key |
| deepSeekAPIKey | string | DeepSeek API key |
| kimiAPIKey | string | Kimi (Moonshot) API key |
| glmAPIKey | string | GLM (Zhipu) API key |
| temperature | number | Sampling temperature |
| maxTokens | number | Maximum output tokens (default: 16384 for Claude, 200000 for other providers). For Claude, values above 21333 automatically enable LangChain streaming to bypass the Anthropic SDK's 10-minute non-streaming guard. |
| agentic | boolean | Enable agentic mode in runPromptWithModel |
| recursionLimit | number | Max agent steps (default: 25) |
| tavilyAPIKey | string | Tavily API key; required when tavily is set |
| tavily | boolean \| TavilyConfig | Enable the Tavily web tools; requires agentic: true (see Web Search (Tavily)) |
Per-model cost constants are read from the same object using flat keys: <model>-input-token-costs, <model>-output-token-costs, <model>-image-input-token-costs, <model>-image-output-token-costs, and <model>-character-costs (all USD per million). Tavily is billed separately via tavily-credit-costs (USD per credit).
License
Not licensed — proprietary software of StackFactor Inc.
Development and release
See docs/README.md — branch flow, how to cut a release, and troubleshooting.
