npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

callscript

v0.1.6

Published

a tool-calling script language for LLMs - the model authors ONE small script (plain JS, compiled to a validated, bounded, inert JSON plan - never executed) instead of calling tools one at a time. Adapter-based: mount AI SDK tools or plain { name, execute

Readme

callscript

Code Mode, without the sandbox.

Experimental. callscript is a Vercel Labs project. APIs may change between releases.

The model writes JavaScript - and it never runs. Instead of calling tools one at a time, the model authors ONE small script in TypeScript and it compiles into an inert JSON plan: validated whole before anything runs (every issue reported at once), executed with bounded call counts, steps scheduled by data dependency. The only things that can execute are the tools you mounted.

Install

npm install callscript

Two small runtime dependencies (acorn, zod), no sandbox, no service: the engine is a library that runs wherever your JS runs. The adapter's peer (ai) is optional and only needed by its own entrypoint.

Quick start

Mount your AI SDK tools on an engine and hand the model the ready-made tools - execute runs the script it authors; search and describe let it discover mounted tools without carrying every definition in its prompt:

import { generateText } from "ai";
import { callscript } from "callscript";
import { toAISDKTools, fromAISDKTools } from "callscript/ai-sdk";

const engine = callscript({
	tools: fromAISDKTools(tools, { namespace: "github" }),
});

await generateText({
	model: "anthropic/claude-sonnet-5",
	prompt: "Close every stale open issue in the 'api' repo.",
	tools: toAISDKTools(engine), // execute + search + describe
});

Why

Classic tool calling has a shape problem: one call per model round-trip. Every intermediate result travels back through the context window just so the model can look at it and decide the next call: slow, token-expensive, and the plan lives nowhere you can inspect it.

Code Mode fixes this by having the model write real TypeScript that calls the tools, then executing that program. The insight is right: models are better at writing programs than at emitting tool-call chains, and intermediate data should flow between calls without a detour through the context. But the price is that you are now executing arbitrary LLM-authored code, so you need a sandbox: an isolate, a container, a worker. That is infrastructure to run, and a security boundary you must get exactly right.

callscript is Code Mode without the sandbox. The model still authors one program - in the JavaScript it already writes, with loops, branches, and dataflow - but nothing it writes ever executes as code: the script compiles into inert JSON with a pure expression subset. The only things that can execute are the tools you mounted. That one change buys the rest:

  • No sandbox needed: there is no arbitrary code to contain. Expressions are a side-effect-free JS subset (no I/O, no globals, no imports), so the engine runs wherever your JS runs: serverless, edge, the same process as your app.
  • Statically checkable: because the plan is data, the whole thing validates before anything runs: unknown tools, misshaped args, unbound references, every issue reported at once. Arbitrary code can only fail at runtime, one error at a time.
  • Bounded by construction: fan-outs declare max and the engine enforces call limits; a runaway loop is a validation error, not an incident.
  • Serializable and resumable: a JSON plan plus a serializable record means runs can suspend (approval gates, external events) and resume with settled steps reused. Pausing arbitrary code mid-execution is famously hard.
  • Parallel for free: steps are scheduled by data dependency, so independent calls run concurrently without the model having to write Promise.all correctly.
  • Plain-value results: sandboxed code hands results back as logged text you parse afterwards; script steps return plain values the next step - and the host - reads directly.
  • Auditable: what will run and what ran are the same inert artifact: reviewable before execution, replayable after.

Expressions are a small language for pure transforms between calls, not general computation. If a step genuinely needs code, then you can have a code tool exposed to the agent.

The JS surface

The authoring format is the JavaScript the model already writes - so the entire language instruction (jsLanguageCard, what engine.describe() puts in the prompt) is one card of ~500 tokens, and the validator's pointed messages teach the rest on retry:

// close stale issues
const issues = await github.listIssues({ repo: "api" });
const stale = issues.filter(i => i.stale);
if (stale.length === 0) return { closed: 0 };
const closed = await Promise.all(
  stale.slice(0, 10).map(i => github.closeIssue({ repo: "api", number: i.number })));
return { count: closed.length };

This is parsed, never executed: each statement compiles into a step of the same inert plan - const x = await tool(...) is a call step, const x = expr a pure derivation, if (cond) return v a guard, Promise.all(list.map(...)) a bounded fan-out (the visible slice is the bound), try/catch the error branch (e compiles to $errors.<id>), and the trailing return the output projection. What is stored, hashed, validated, approved, and resumed is always the JSON plan; the JS text is a frontend, not a runtime.

Semantics match the JS reading: awaited calls run in statement order (the compiler inserts after edges where no data already orders them), and Promise.all - the spelling the model already knows for concurrency - runs calls in parallel. A call's second argument carries per-step options: github.closeIssue(args, { reason: "why", suspend: true, onError: "skip" }). An un-awaited call to a mounted tool detaches (fire-and-forget) and a later script joins it with const r = await job.

Everything outside the recognized grammar is rejected before anything runs, with the message that names the callscript spelling - while points at the bounded fan-out, reassignment at single-assignment consts, new Set at the dedupe idiom - so one retry converges:

while (queue.length) { await github.closeIssue({ number: next }); }
  ✗ line 1: unbounded loops cannot run - fan out over a bounded list
    instead: await Promise.all(items.slice(0, N).map(item => tool.name({ ... })))

closed = await github.closeIssue({ number: 7 });
  ✗ line 4: bindings are single-assignment - declare a new const instead of reassigning

const seen = new Set(ids);
  ✗ line 7: Unsupported syntax: new. Dedupe with
    xs.filter((x, i, a) => a.indexOf(x) === i); group with Object.groupBy(xs, x => x.key).

Every door accepts both surfaces: a string is the JS surface, an object is the JSON plan. Engines default to teaching the JS surface (describe, toolDefinition, tools); pass format: "json" to teach the JSON plan instead - execution accepts both regardless.

The plan: three verbs, four modifiers

Under the surface, a script is one inert JSON plan: a list of steps wired by dataflow. Steps reference each other by id, and those references ARE the schedule - independent steps run concurrently, dependent ones wait for their inputs. Each step is one of three verbs:

  • call - const x = await tool.name({...}) - invokes a mounted tool; its args validate against the tool's schema before it fires.
  • let - const x = expr - derives a value from earlier steps with a pure expression.
  • return - if (cond) return value - a guard clause: when it fires the run ends right there with that value; otherwise the run continues.

And a step can carry modifiers:

  • if skips the step unless a condition holds.
  • each fans a call out over a list, one dispatch per element, bounded by a hard max.
  • after orders a step behind earlier ones when no data flows between them - close the issues, then post the summary.
  • suspend flags a call for confirmation: the run pauses there until a human approves it.

A few more things the language gives you:

  • Globals. Expressions read earlier steps by id, input (data passed to this execution), variables published by earlier runs in the session, $errors.<stepId> for recorded failures, and safe built-ins like Math, JSON, and Date. new Date(...) gives an ISO 8601 string; the read-only Date API (getUTCDate(), getTime(), toISOString(), ...) works on it.
  • Promises. Every call is async; await only decides whether the run blocks on it. A call without await detaches and keeps running in the background; a later script joins it with const r = await job.
  • Expressions. A side-effect-free subset of JS: arrows, template literals, ternaries, optional chaining - no I/O, no imports, no reaching outside the script's scope.
  • Output. output projects the run's final result from any settled step; by default it is the last step's value.

Limits

Every engine carries hard limits: validation enforces them before a run starts, execution enforces them while it runs, and engine.describe() renders the live numbers into the prompt so the model authors against the same bounds the engine enforces.

| limit | default | bounds | | --- | --- | --- | | maxSteps | 20 | steps per script | | maxItemsPerStep | 100 | the max any single each fan-out may declare | | maxTotalCalls | 200 | worst-case total calls per script (Σ max) | | maxConcurrency | 5 | independent steps / fan-out calls in flight | | maxExprNodes | 100 000 | AST nodes evaluated per expression | | maxCallResultBytes | 10 MiB | serialized size of a single call's result | | maxSuspendAttempts | 5 | times one suspension key may re-raise before failing |

Override any subset: callscript({ tools, limits: { maxTotalCalls: 50 } }).

Adapters

The engine is adapter-based. It never knows where a tool came from; everything mounts through one neutral shape, ScriptTool:

{ name, description?, inputSchema?, outputSchema?, errors?, execute(args, ctx) }

so you can use callscript purely with the AI SDK, with plain object literals, or any mix:

  • callscript/ai-sdk: hand it the same tools record you'd give generateText/streamText
  • plain literals: a { name, execute } object is already a tool; the tool() helper pins the literals for typed authoring

With the AI SDK

AI SDK in, AI SDK out: you never hand-write a script. Define tools with the SDK's own tool(), mount them through the adapter, and toAISDKTools(engine) hands the model the engine as ready-made tools: execute runs the script it authors, search and describe discover the mounted tools on demand.

import { generateText, tool } from "ai";
import { z } from "zod";
import { callscript } from "callscript";
import { toAISDKTools, fromAISDKTools } from "callscript/ai-sdk";

// the same `tool()` objects you'd hand generateText directly
const listIssues = tool({
	description: "list the issues of a repo",
	inputSchema: z.object({ repo: z.string() }),
	execute: async ({ repo }) => [/* ... */],
});
const closeIssue = tool({
	description: "close an issue by number",
	inputSchema: z.object({ repo: z.string(), number: z.number() }),
	execute: async ({ number }) => ({ closed: number }),
});

const engine = callscript({
	// the whole record mounts at once, however many tools it holds;
	// `namespace` prefixes the keys: github.listIssues, github.closeIssue
	tools: fromAISDKTools({ listIssues, closeIssue }, { namespace: "github" }),
});

await generateText({
	model: "anthropic/claude-sonnet-5",
	system: "You act by writing ONE callscript.",
	prompt: "Close every stale open issue in the 'api' repo.",
	tools: toAISDKTools(engine), // `execute` runs a script; `search`/`describe` find tools
});

toAISDKTools(engine) wraps engine.tools(), the host-neutral trio. execute is the tool: it validates the script at the door (a rejected script returns status: "invalid" with every issue, for the model to retry), runs it against the pair's accumulated session, and returns the outcome without the session state (that stays host-side, never in the prompt). search and describe exist so the prompt doesn't carry every tool definition ahead of time: search matches mounted tools by keyword and returns names with one-line summaries, describe returns the full signature cards for the names a script will actually call. With 20 or fewer tools the cards inline into execute's description; past that (or with inlineTools: false) the model discovers tools through search/describe instead, so the prompt stays small however many tools you mount. Pass state to seed the session and onState to observe the record after every run (persist it anywhere - it is plain data), names to rename the tools. Prefer manual wiring? engine.toolDefinition() still returns the one-tool description/inputSchema to hand a tool interface yourself.

For the task above the model authors one script - which the engine compiles into the inert plan, no code ever running:

// close stale issues
const issues = await github.listIssues({ repo: "api" });
const stale = issues.filter(i => i.stale);
const closed = await Promise.all(
  stale.slice(0, 10).map(i => github.closeIssue({ repo: "api", number: i.number })));

which is stored and executed as:

{
	"intent": "close stale issues",
	"steps": [
		{ "id": "issues", "call": "github.listIssues", "args": { "repo": "api" } },
		{ "id": "stale", "let": "issues.filter(i => i.stale)" },
		{
			"id": "closed",
			"call": "github.closeIssue",
			"each": "(stale.slice(0, 10)).map((i) => ({ repo: \"api\", number: i.number }))",
			"max": 10
		}
	]
}

The record keys are the registry: a script's call names a tool by key, and validate rejects anything else before a run starts. A namespace prefixes the keys, so whole records mount side by side by spreading: tools: [...fromAISDKTools(github, { namespace: "github" }), ...fromAISDKTools(slack, { namespace: "slack" })]. Args validate against each tool's schema (zod, any standard schema, or jsonSchema()) before execute fires; a failure fails the step with code invalid_tool_args. The schemas also render the tool cards engine.describe() / engine.toolDefinition() put in the model's prompt.

With plain tools

import { callscript, tool } from "callscript";

const closeIssue = tool({
	name: "github.closeIssue",
	description: "close an issue by number",
	inputSchema: { type: "object", properties: { repo: { type: "string" }, number: { type: "number" } }, required: ["repo", "number"] },
	errors: ["not_found"],
	execute: (args: { repo: string; number: number }) => ({ closed: args.number }),
});

const engine = callscript({ tools: [closeIssue] });

Throw protocol from execute: throw to fail the step (a string code on the error surfaces as the step error's code), throw earlyReturn(value) to end the run here, throw suspend({ key, ... }) to park it on an external event.

Runs, suspensions, per-execute input

engine.run validates at the door, executes with the engine's limits, and returns the ordinary ExecuteResult, with early returns, suspensions, and the serializable session state included. Per-execute data flows through expressions:

// first run suspends; re-execute with the answer piped in as `input`
const first = await engine.run({ script }); // status: "suspended"
const second = await engine.run({ script, state: first.state, input: { code: "42" } });

The suspended run comes back as a serializable state record: store it anywhere - it survives a restart - and settled steps are reused instead of re-run on resume.

Sessions

A session is just the accumulated run record - result.state, a plain serializable value. Thread it to continue: settled steps are reused, and a later script references an earlier run's step outputs by name - no re-fetch:

const one = await engine.run({ script: first });
const two = await engine.run({ script: second, state: one.state }); // reads `first`'s step outputs

Keep it in a variable, or JSON.stringify it into any KV/database between turns - it survives a restart. For long-lived suspended runs, the durable runner persists the whole lifecycle behind a four-call storage contract.

The error branch of the dataflow

A step that declares "onError": "skip" records its failure instead of failing the run, and later expressions consume it as $errors.<stepId>: UCAN's await/error, spelled with a step id. $errors.close is the { message, code } of step close when it failed (a per-element list on an each fan-out), undefined when it succeeded, and the read creates the same dependency edge a value reference would. In the JS surface it is just try/catch:

try {
  const closed = await github.closeIssue({ number: 42 });
} catch (e) {
  await slack.post({ text: `close failed: ${e.message}` });
}

which compiles to:

{ "id": "closed", "call": "github.closeIssue", "args": {}, "onError": "skip" },
{ "if": "$errors.closed", "call": "slack.post",
  "args": { "text": "=`close failed: ${$errors.closed.message}`" } }

Scripts compile into tools

engine.tool(name, script, { description? }) turns a script into a ScriptTool, mountable on another engine, so a script becomes a tool of another engine, and the signals compose: an inner approval gate ends the hosting run early, an inner suspension parks it, and the answer flows back down through args (args: { approved: "=input.approved" }). Free names no step produces are collected as external, the script's session requires, checked against the live session on every call.

const closeStale = engine.tool("github.closeStale", script);

let state;
const persist = (s) => (state = s); // hands the settled record back, even when a gate throws
await closeStale.execute({}, { state, persist });                 // gate → throws EarlyReturnSignal({ confirm: [...] })
await closeStale.execute({ approved: true }, { state, persist }); // resumes - settled steps reused

Typed authoring

engine.script({...}) and engine.tool(...) take a TypedScript of the engine: call is the union of mounted tool names, args is that tool's argument type (the first parameter of its execute) with any leaf (or subtree) replaceable by an expression, so a typo'd tool name or a misshaped arg is a type error before it is a validation error. Adapter output carries the types through, so this works over fromAISDKTools(...) too.

Every expression position takes the string form or a real JS arrow, transpiled (never executed) into the string at the door:

engine.script({
	steps: [
		{ id: "issues", call: "github.listIssues", args: { repo: "api" } },
		{ id: "stale", let: ({ issues }) => issues.filter((i) => i.stale) },
		{
			call: "github.closeIssue",
			each: ({ stale }) => stale.map((issue) => ({ repo: "api", number: issue.number })),
			max: 10,
		},
	],
});

The arrow's parameter is the scope contract: a plain destructuring naming everything the body reads (input and $calls included; aliases allowed; the key is the env name). The body must be a single expression in the same grammar as the strings, and a free name, including a captured outer variable (the thing a native closure could otherwise smuggle in), is rejected at the door. What's stored, hashed, rendered, and re-executed is always the string form, so the script stays inert data.

The language core (the pure-JS expression subset, static validation, limits, records, reconciliation, the runner) is engine-independent - parseJsScript and validateScript are usable standalone, and the engine is the door onto the rest.

Examples

Runnable tours, no build step:

bun examples/t.ts       # describe/toolDefinition/context - the three prompt layers
bun test/ai-sdk.ts      # end to end with the AI SDK: model authors, engine validates + runs

License

Apache-2.0