@stackfactor/agent-utils
v1.9.0
Published
Shared utilities for StackFactor AI agent services — LangChain helpers, structured logging, error handling, and constants.
Keywords
Readme
@stackfactor/agent-utils
Shared utilities for StackFactor AI agent services — LangChain helpers, structured logging, error handling, and constants.
Installation
npm install @stackfactor/agent-utilsModules
The package exports four modules:
import {
langChain,
logger,
errorHandling,
constants,
} from "@stackfactor/agent-utils";langChain
Unified interface for running LLM prompts, managing LangChain agents, and generating images across the OpenAI, Anthropic, and Google providers.
langChain.runPromptWithModel(modelName, config, prompt, onProgressReport?, minPercent?, maxPercent?, expectsJsonResponse?, schema?, agentName?, tools?, usageTracker?)
Sends a prompt to an LLM and returns the response. Supports three execution modes:
- Agentic (
config.agentic === true) — creates and runs a LangChain agent with tools. - Streaming (
onProgressReportprovided) — streams the response with progress callbacks. Uses nativeresponse_formatfor OpenAI+schema, or NDJSON for other providers. - Non-streaming — direct invocation with JSON extraction.
When expectsJsonResponse is true (the default), JSON escape instructions are injected and the response is parsed. If a Zod schema is provided, the parsed result is validated.
| Parameter | Type | Default | Description |
| --------------------- | -------------------- | --------------- | -------------------------------------------------------------------------------------------------------- |
| modelName | string | — | Model identifier: "gpt-4o", "claude-3-5-sonnet", "gemini-1.5-pro", etc. |
| config | object | — | API keys (openAIAPIKey, anthropicAPIKey, googleAPIKey), temperature, agentic, recursionLimit, tavily, tavilyAPIKey |
| prompt | string \| object[] | — | Plain string or array of { role, content } message objects |
| onProgressReport | function \| null | null | Async callback receiving { message, progress } updates |
| minPercent | number | 0 | Lower bound for progress percentage |
| maxPercent | number | 100 | Upper bound for progress percentage |
| expectsJsonResponse | boolean | true | Parse response as JSON |
| schema | ZodSchema \| null | null | Zod schema for response validation |
| agentName | string | "StackFactor" | Display name for the agent (agentic mode) |
| tools | any[] | [] | LangChain tools available to the agent (agentic mode) |
| usageTracker | UsageTracker \| null | null | Optional counter, for the agent's own use, of what each model used. The platform does not read it (see Usage). |
Returns: JSON string (when expectsJsonResponse is true) or raw content string.
const usageTracker = {};
const result = await langChain.runPromptWithModel(
"gpt-4o",
{ openAIAPIKey: "sk-..." },
"Generate a summary of this document.",
null,
0,
100,
true,
null,
"StackFactor",
[],
usageTracker,
);
// usageTracker is now:
// {
// cost: 0, // always: the platform prices what was used
// tokens: {
// "gpt-4o_inputTokens": 412,
// "gpt-4o_outputTokens": 87,
// }
// }langChain.createAIAgent(name, modelName, systemPrompt, tools?, responseFormat?, config, onReportProgress?, minPercent?, maxPercent?)
Constructs a LangChain agent with a model, system prompt, and tools. When onReportProgress is provided, a report_progress tool is automatically added.
| Parameter | Type | Default | Description |
| ------------------ | ------------------ | ------- | ---------------------------------------- |
| name | string | — | Display name for the agent |
| modelName | string | — | LLM identifier |
| systemPrompt | string | — | System prompt describing agent behaviour |
| tools | any[] | [] | LangChain tool instances |
| responseFormat | any | — | Structured response format descriptor |
| config | object | — | API keys, temperature, tavily, etc. |
| usageTracker | object \| null | null | Accumulator the Tavily tools bill credits and record sources into |
| onReportProgress | Function \| null | null | Progress callback |
| minPercent | number | 0 | Minimum reportable progress |
| maxPercent | number | 100 | Maximum reportable progress |
Returns: A configured LangChain agent instance.
langChain.runAIAgent(agent, prompt, config, onProgress?)
Executes a LangChain agent with a user prompt. Registers callbacks for tool_start, tool_end, agent_action, and error events when onProgress is provided. Logs execution time on completion.
| Parameter | Type | Default | Description |
| ------------ | ------------------ | ------- | -------------------------------- |
| agent | any | — | Agent created by createAIAgent |
| prompt | string | — | User message to send |
| config | object | — | recursionLimit (default: 25) |
| onProgress | Function \| null | null | Progress event callback |
Returns: Raw response from the agent's invoke method.
const agent = langChain.createAIAgent(
"Summarizer",
"gpt-4o",
"You summarize documents concisely.",
[],
null,
{ openAIAPIKey: "sk-..." },
);
const response = await langChain.runAIAgent(
agent,
"Summarize this text...",
{},
);langChain.runPromptWithModelForImageGeneration(modelName, config, prompt, options?)
Generates an image using OpenAI or Google AI models.
| Parameter | Type | Default | Description |
| ----------- | -------- | ------- | --------------------------------------------------------------------------------------------------------- |
| modelName | string | — | Image model: "dall-e-3", "gpt-image-1.5", "imagen-4.0-generate-001", "gemini-3.0-pro-image", etc. |
| config | object | — | openAIAPIKey or googleAPIKey |
| prompt | string | — | Text prompt describing the image |
| options | object | {} | Provider-specific options (see below) |
OpenAI options: size ("1024x1024", "1792x1024", etc.), style ("vivid" or "natural"), responseFormat ("url" or "b64_json"), n (number of images).
Google options: aspectRatio ("1:1", "3:4", "4:3", "9:16", "16:9"), numberOfImages, negativePrompt.
Returns: { url?, b64_json?, revisedPrompt? } for single images, or { images: [...] } for multiple.
const image = await langChain.runPromptWithModelForImageGeneration(
"dall-e-3",
{ openAIAPIKey: "sk-..." },
"A futuristic city skyline at sunset",
{ size: "1792x1024" },
);langChain.throwErrorIfNotSuccessful(response)
Guards that a response is a non-empty string. Throws an INTERNAL_SERVER_ERROR if not.
| Parameter | Type | Description |
| ---------- | ----- | ----------------- |
| response | any | Value to validate |
Returns: The response string if valid.
logger
GCP-compatible structured logging via Winston with OpenTelemetry trace enrichment.
logger.log(request, level, message, options?)
Writes a structured log entry. Automatically enriches with OpenTelemetry traceId and spanId when an active span exists. Prepends the user's email from the request object when available.
| Parameter | Type | Default | Description |
| --------- | ---------- | ------- | -------------------------------------------------------------------------- |
| request | any | — | HTTP request object (reads request.user.email), or null |
| level | LogLevel | — | "error", "warn", "info", "http", "verbose", "debug", "silly" |
| message | string | — | Log message |
| options | object | {} | Additional structured fields (service, requestId, etc.) |
logger.log(req, logger.levels.info, "User signed in", { service: "auth" });
logger.log(null, logger.levels.error, "Connection failed");logger.levels
Enum-like object mapping level names to their string values:
logger.levels.error; // "error"
logger.levels.warn; // "warn"
logger.levels.info; // "info"
logger.levels.http; // "http"
logger.levels.verbose; // "verbose"
logger.levels.debug; // "debug"
logger.levels.silly; // "silly"errorHandling
Factory for consistently shaped error payloads.
errorHandling.create(errorCode, errorMessage, details?)
Creates a structured error object with a stack trace.
| Parameter | Type | Default | Description |
| -------------- | -------- | ------- | ------------------------------------------ |
| errorCode | number | — | HTTP status code or application error code |
| errorMessage | string | — | Human-readable error description |
| details | any | null | Optional supplementary information |
Returns: { code, message, details, stack }
throw errorHandling.create(404, "User not found");
throw errorHandling.create(400, "Validation failed", { field: "email" });constants
Shared constants used across the package.
constants.HTTP_CODES
| Constant | Value |
| ----------------------- | ----- |
| BAD_REQUEST | 400 |
| UNPROCESSABLE_ENTITY | 422 |
| INTERNAL_SERVER_ERROR | 500 |
| BAD_GATEWAY | 502 |
constants.ERROR
| Constant | Value |
| ---------------------------- | ---------------------------------------- |
| UNABLE_TO_GENERATE_CONTENT | "Unable to generate content" |
| UNEXPECTED_ERROR | "An unexpected error occured..." |
| UNSUPPORTED_MODEL | "The specified model is not supported" |
platform
Calls an integration makes back to StackFactor on its own behalf.
platform.getTenantSettings(token)
For scheduled runs (event === 510). A scheduled run starts once for the whole platform and
carries no identity: request.authToken is absent and session is null. The integration
brings its own: create a service account and an API token in the admin tenant, store the token as
a secret constant of the integration, and pass it here. The answer is this integration's settings
in every organization that has it enabled — an organization that switched it off is not listed.
const { platform } = require("@stackfactor/agent-utils");
// params.data.schedule → { name, cron, timezone, scheduledTime }
const tenants = await platform.getTenantSettings(config.STACKFACTOR_API_TOKEN);
for (const { tenant, settings } of tenants) {
// settings: the tenant-level fields this integration declares, secrets included.
// To act INSIDE the tenant, declare a secret field for the tenant's own API key
// and use settings.<thatField> for those calls.
}Which integration is asking is never yours to say: the platform puts a signed run ticket in the
run's envelope (request.runTicket) and the helper sends it back with your token. You never touch
it, and it cannot be used to read another integration's settings.
The run lifecycle (job executions)
An agent subscribed to a platform event with runAs: job is started as one Cloud Run job
execution per event. The runtime reports that run to the platform on the handler's behalf, and
the handler neither does nor knows any of it:
| When | What happens | Why it matters to you |
|---|---|---|
| Before the handler | Claim. Of however many executions the platform may start for one run — a retried task, a redelivered message — exactly one wins. | An execution that loses the claim exits 0 without running the handler: the work is somebody else's, and exiting 1 would have Cloud Run retry it forever. Write handlers as if they run once; they do. |
| While it runs | Heartbeat, every 30s. | A run that stops beating is treated as abandoned and swept. A handler that blocks the event loop for minutes will look dead — do long work in awaited chunks. |
| After the handler | Settle, COMPLETED with whatever the handler returned, or FAILED with the error. | The run's record carries your return value as its outcome, so return something small and descriptive. Settling also revokes the execution's token, so do not hold it past your handler. |
Which run is being reported is never the execution's to assert: the platform signs a run ticket
into the envelope (request.runTicket) and the runtime presents it. A run whose envelope carries
no ticket is not claimed or settled — nothing names it.
A settle the platform refuses (settled: false — the run was already swept) or that fails
outright does not change the execution's outcome: the work either happened or it did not.
Usage
What an agent used reaches the platform one way: the usage report. What a handler returns is its content and nothing else. A handler that returns [data, usage] sends that pair to the caller as the content.
Every model call made through langChain and every Tavily call is counted for the request it ran in. serve and runOnce report what was counted once a minute while the handler runs, once more when it ends, and also when it throws. The agent prices nothing: it reports a model, a unit and an amount, and the platform applies its own rates and charges the tenant's AI credits.
An agent that calls a provider without langChain counts the call itself:
import { recordUsage } from "@stackfactor/agent-utils";
const response = await anthropic.messages.create({ model, messages });
recordUsage(model, "inputTokens", response.usage.input_tokens);
recordUsage(model, "outputTokens", response.usage.output_tokens);| Unit | Counts |
| ------------------- | ------------------------------- |
| inputTokens | Text tokens sent to a model |
| outputTokens | Text tokens a model answered |
| imageInputTokens | Image tokens sent to a model |
| imageOutputTokens | Image tokens a model generated |
| characters | Characters turned into speech |
| seconds | Seconds of generated video |
| credits | Credits of a metered service |
recordUsage counts nothing outside a request, and nothing for an amount that is not a positive number. A report that fails is sent again twice under the same id, so it is charged once; one that still fails is logged and the run goes on.
Web Search (Tavily)
Set config.tavily to give an agent the Tavily web tools. These run client-side: the model emits a tool call, this library performs the HTTP request, and the result is appended to the conversation. Only an agent loop can execute them, so config.agentic must be true — enabling tavily without it throws BAD_REQUEST rather than handing the model tools whose calls nothing answers.
| Tool | Arguments | Returns | Default |
| ------------- | -------------------------------- | ----------------------------------------------------------------------- | ------- |
| web_search | query, maxResults? | Snippets and URLs; no page content | on |
| web_extract | urls[], query? | Page content — matching passages with a query, whole document without | on |
| web_map | url, instructions? | URLs only, no content | off |
| web_crawl | url, instructions?, limit? | Full content of every page visited | off |
No new dependency is added: all four endpoints are a single POST, issued with the runtime's own fetch.
Both reading modes, one tool
web_extract covers whole-page and targeted reading through Tavily's own optional query field, so the model chooses per call:
queryomitted → the entire document.querysupplied → only the matching chunks (chunks_per_source, default3).
Targeted extraction is dramatically cheaper in context and is usually sufficient; whole-document extraction is there for when the full structure matters. Splitting these into two tools would duplicate the description tokens on every request to express a distinction the API already makes with one optional field.
const usageTracker = {};
const result = await langChain.runPromptWithModel(
"claude-opus-5",
{
anthropicAPIKey: "sk-ant-...",
tavilyAPIKey: "tvly-...",
agentic: true, // required
tavily: {
tools: ["search", "extract", "map"],
includeDomains: ["sc.gov", "sc.edu"],
maxCharsPerUrl: 60000,
},
},
"What are the SC procurement thresholds for sole-source awards?",
null, 0, 100, false, null, "StackFactor", [], usageTracker,
);
// usageTracker.tokens.tavilyCredits === 4
// usageTracker.webSearchSources === [{ url: "https://...", title: "..." }, ...]TavilyConfig
tavily: true uses the defaults below. Pass an object to configure it.
| Field | Default | Description |
| ----------------- | ------------------------ | -------------------------------------------------------------------------------------------- |
| tools | ["search", "extract"] | Which tools to expose; each one spends its description tokens on every request |
| searchDepth | Tavily's "basic" | "basic" \| "advanced" \| "fast" \| "ultra-fast" |
| extractDepth | Tavily's "basic" | "advanced" also pulls tables and embedded content, at double the credits |
| maxResults | 5 | Results per search |
| chunksPerSource | 3 | Relevant chunks per URL when a query is passed to web_extract |
| maxCharsPerUrl | 60000 | Per-URL cap before truncation (≈15k tokens) |
| maxCharsTotal | 150000 | Cap across one tool call (≈37k tokens) |
| includeDomains | — | Restrict web_search to these domains |
| excludeDomains | — | Exclude these domains from web_search |
| crawlLimit | 10 | Maximum pages one web_crawl or web_map call may return |
| format | "markdown" | "markdown" preserves table and heading structure at fewer tokens than the prose equivalent |
Context management
Nothing sits between Tavily's response and the context window — there is no provider-side filtering here — so the caps above are the only thing standing between one extract call and a 100-page PDF. Three measures apply by default:
Responses are stripped before the model sees them. Tavily returns score, published_date, favicon and id alongside each result. None of it informs an answer, so only the title, URL and text are forwarded; the rest is dropped rather than billed as input tokens.
web_search never requests page content. include_raw_content is always false: search finds documents, web_extract reads them. Asking for full text at search time pays for every result to find one.
Truncation is reported, never silent. When a document exceeds maxCharsPerUrl, or a call exhausts maxCharsTotal, the payload says exactly what was cut:
[truncated: 60,000 of 412,336 characters shown — narrow the query or request fewer URLs to see the parts that matter]A model that cannot tell it received half a regulation will reason over the half it got, which on a compliance corpus is worse than an error.
web_crawl is off by default for the same reason: it returns full content for every page it visits, and is the easiest way here to exhaust a context window. Prefer web_map to enumerate a site, then web_extract with a query on the few URLs that matter.
Notes
- Works with every supported provider.
- Credits come from each response's own
usageblock rather than an estimate. They are reported to the platform ascreditsof the modeltavily(see Usage) and counted onusageTracker.tokens.tavilyCredits. - URLs are collected on
usageTracker.webSearchSources, deduplicated across the whole run — Tavily's terms require citing the original sources when its content is shown to end users. - Per-URL failures are surfaced in the tool result instead of being dropped, so a fetch failure is not mistaken for an absence of content.
- Tool calls honour the caller's abort signal, so a cancelled run stops paying for in-flight crawls.
Checking that cited links resolve
checkUrlsReachable(urls, config, usageTracker?) returns the URLs, exactly as passed, that Tavily could fetch. Use it in place of requesting a link yourself. Cited links come from web pages anyone can publish, so a link, its redirects or its DNS can point inside StackFactor's network; a request this process sends would go there through StackFactor's own egress addresses. Tavily fetches from its own infrastructure.
const { checkUrlsReachable } = require("@stackfactor/agent-utils");
const live = await checkUrlsReachable(references.map((r) => r.url), config, usageTracker);
const kept = references.filter((r) => !r.url || live.has(r.url));- Only
http:andhttps:URLs are sent, 20 per call (Tavily's limit), at basic depth: 1 credit per 5 URLs fetched, failed ones free. Credits are billed like any Tavily call; the URLs are not added towebSearchSources. - A URL Tavily lists under
failed_resultsis unreachable. A batch whose call fails, or a config withouttavilyAPIKey, confirms none of its URLs. It never throws. - No
config.tavilyoragenticis needed: it calls Tavily directly, not through the model.
Supported Models
Text / Chat
| Provider | Model Prefix | Example |
| ----------------- | ---------------------- | ------------------------------------ |
| OpenAI | gpt- | gpt-4o, gpt-4o-mini |
| Anthropic | claude- | claude-3-5-sonnet, claude-3-opus |
| Google | gemini- | gemini-1.5-pro, gemini-2.0-flash |
Native structured output is configured per provider when a Zod schema is passed: response_format with a JSON schema for gpt-*, output_config.format for claude-*, and JSON mode plus responseSchema for gemini-*.
Output token limits are not configurable. Anthropic requires max_tokens on every request, so claude-* is pinned to 64000 with streaming always enabled; gpt-* and gemini-* are sent no limit and use each provider's own default, which is that model's maximum.
Image Generation
| Provider | Model Prefix | Example |
| -------- | ----------------------- | ------------------------------------------------- |
| OpenAI | dall-e-, gpt-image- | dall-e-3, gpt-image-1.5 |
| Google | imagen-, gemini- | imagen-4.0-generate-001, gemini-3.0-pro-image |
Config Object
The config object accepted by LangChain methods supports the following keys:
| Key | Type | Description |
| ----------------- | --------- | ------------------------------------------- |
| openAIAPIKey | string | OpenAI API key |
| anthropicAPIKey | string | Anthropic API key |
| googleAPIKey | string | Google AI API key |
| temperature | number | Sampling temperature |
| agentic | boolean | Enable agentic mode in runPromptWithModel |
| recursionLimit | number | Max agent steps (default: 25) |
| tavilyAPIKey | string | Tavily API key; required when tavily is set |
| tavily | boolean \| TavilyConfig | Enable the Tavily web tools; requires agentic: true (see Web Search (Tavily)) |
The object carries no prices. The <model>-…-costs and tavily-credit-costs constants are no longer read: the platform prices what an agent used (see Usage).
License
Not licensed — proprietary software of StackFactor Inc.
Development and release
See docs/README.md — branch flow, how to cut a release, and troubleshooting.
