@upship/agents
v0.2.12
Published
An agent platform where agents are data: versioned rows with a prompt, a model tier, tools, and an optional pipeline. API, CLI, and web UI in one install.
Downloads
2,427
Readme
Agent Platform
Docs and project site: https://khtdr.com/agents/
Agents are data, not code. An agent is a Postgres row (or SQLite row, or JSON file) with a system prompt, a tier, and a list of tool names — created, edited, versioned, and run through an HTTP API, a CLI, or a web UI. Nothing about adding an agent requires a rebuild.
Tools are the other half: those are code, registered at startup, with their
descriptions and runtime config stored as data rows layered on top. And since an
agent is addressable by name just like a tool, one agent can list another in its
tools array — orchestration falls out of the same lookup.
Install
npm install -g @upship/agents
agent serve # asks for your OpenRouter key once, then API + UI on :2137Everything an install owns lives under ~/.agents (AGENTS_HOME to move it):
its .env, a SQLite database, and the directories tools read from and write
to. Agents read the install's data directory and the places your projects live
(DATA_DIR in that .env, a list; first run fills in ~/code and friends if
they exist, and agent add-data-dir . adds where you stand), so cd into a
project and run an agent against it. Then, from any
other terminal:
agent new # it asks what you want, then drafts an agent
agent list
agent <name> "some input"Or run the server in the background and manage it from anywhere:
agent start # same server, detached; waits until it answers
agent open # the web UI in a browser
agent status # up? where? since when?
agent logs -f # what it has been doing
agent stopNeeds Node 24 or newer. Everything below is for working on the platform itself.
Quick start (from a checkout)
npm install
cp .env.example .env # fill in OPENROUTER_API_KEY
npm run db:push # set up whatever STORE_DRIVER points at
npm run dev # API on :2137
npm run dev:ui # UI on :2139The shortest useful .env is two lines:
OPENROUTER_API_KEY=sk-or-...
STORE_DRIVER=sqliteSTORE_DRIVER=sqlite needs no database to stand up (Node 24+, which is where
node:sqlite stabilized). Set it in .env rather than per-command — the server,
the CLI, and the seed script each resolve it independently, so a value typed on
one command line configures only that process.
Seed some example agents with npm run seed. To run both processes under pm2
instead of two terminals, npm run pm2:start.
Then make your first agent — from here or from anywhere else:
agent new # it asks what you want, then drafts one
agent skill install # or teach Claude Code to author here: /new-agent
agent brief # or just read what this install can doOr make the runtime part of your own app rather than a service beside it:
cd ~/my-app
agent bundle && npm install ./lib/agents # vendors the runtime; see lib/agents/README.mdConcepts
An agent
{
"name": "researcher",
"systemPrompt": "You research topics and report findings with sources.",
"model": "default",
"tools": ["web_search", "write_file"]
}Run it:
agent researcher "what changed in HTTP/3 this year"tools entries resolve through three tiers, in order: a DB tool definition →
the ToolRegistry → an agent name. That last tier is what makes an
orchestrator just another agent.
Tiers, not models
Agents name a tier — default, fast, or reasoning — never a model id. What a
tier points at is data: shipped defaults, then MODEL_DEFAULT/MODEL_FAST/
MODEL_REASONING env vars, then a stored binding written from the /models
settings page. Repoint a tier and every agent using it moves, with no redeploy.
Prices ride along with the binding, because a model id and its price have to
agree — a mismatch is how a run silently reads as free. An agent that must not
move sets model to an explicit id, which wins over the tier.
Execution modes
Which optional fields an agent sets decides how it runs. Three of them pick an execution mode and the rest wrap whichever one is in play:
| Field set | Behavior |
|---|---|
| (none) | Freeform LLM loop — the model picks tools and decides when it's done |
| graph | Deterministic DAG — nodes execute in topological order, no orchestration tokens |
| planning | The model writes its own graph first, optionally for a person to approve, then runs it |
| debate | N candidates from this agent or a named panel, and a verdict by vote, weight, or judge |
graph and planning are mutually exclusive — a planning agent's graph is the
plan it writes — and a 400 on the save says so. The three wrappers apply to
whichever mode ran:
| Field set | Behavior |
|---|---|
| outputSchema | The final message is parsed and validated against a JSON Schema |
| reflect | Produce → critique → refine, until a critic agent signs off |
| review | A critic reads the finished output once. Nothing re-runs — a run it doesn't approve ends needs_review |
Use a graph when the sequence is known in advance and you want lower cost and per-node error attribution. Use freeform when the model needs to make judgment calls mid-flow.
Graphs
Nodes reference tools or agents by name; edges define the topology. Independent
branches run concurrently, fan-in merges its parents' outputs, and an edge can
carry a condition that routes on the source node's output:
{
"graph": {
"nodes": [
{ "id": "classify", "ref": "intent-classifier" },
{ "id": "web", "ref": "web_search", "input": { "query": "{{input.query}}" } },
{ "id": "docs", "ref": "docs-agent" },
{ "id": "answer", "ref": "summarizer", "merge": "concat" }
],
"edges": [
{ "from": "classify", "to": "web", "condition": "output.intent == 'web'" },
{ "from": "classify", "to": "docs", "condition": "output.intent == 'docs'" },
{ "from": "web", "to": "answer" },
{ "from": "docs", "to": "answer" }
]
}
}Conditions are a closed whitelist grammar parsed by hand in
src/domain/condition.ts — not eval, and not a JS subset that happens to
parse. They arrive as data over HTTP, so "no assignment, no arbitrary calls, no
reachable globals" has to be a property of the parser. The grammar is validated
at save time, so a typo is a 400 rather than a run that dies when it first
reaches that edge.
A node whose every inbound edge resolved without firing is skipped, and skipping prunes its outgoing edges in turn — that cascade is what keeps the fan-in node above from waiting forever on a branch that will never run.
Content blocks
Inputs and outputs are arrays of typed blocks — text, image, file, data.
The data block carries a machine-readable value as itself rather than as JSON
smuggled through text, optionally tagged with the schema it validated against.
When blocks are mixed, {{input}} follows the data block if there is one.
CLI
bin/agent talks to the API over HTTP. Symlink it onto your PATH:
ln -s "$PWD/bin/agent" ~/bin/agentagent list # available agents
agent run researcher "topic" # or just: agent researcher "topic"
agent researcher -f notes.md -f chart.png # attach files and images
echo "topic" | agent run researcher # stdin
agent prompt -t web_search -t write_file "…" # ad-hoc run, no stored agent
agent pending # runs waiting on a person, answerable here
agent new # asks what you want, drafts it, shows it, writes it
agent skill install # write the /new-agent skill for Claude Code
agent brief # this install's authoring brief
agent bundle # vendor the runtime into the project you are inagent new, agent skill install and agent brief are the way in that does not
need a checkout: the brief is generated from the live registries and served at
GET /api/brief, so it lists the tools this install really has, and both the
skill and the drafting agent read that one document rather than a copy.
It detects a TTY: interactive sessions get streaming, spinners, and box drawing;
piped output gets plain text. AGENT_API_URL points it at a remote deployment.
API
Base URL http://localhost:2137.
| Method | Path | |
|---|---|---|
| POST / GET | /api/agents | create / list |
| GET / PUT / DELETE | /api/agents/:name | get / update / soft delete |
| POST | /api/agents/:name/run | run synchronously |
| POST | /api/agents/:name/stream | run with streaming |
| GET | /api/agents/:name/versions | version history |
| POST | /api/agents/:name/replay | re-run with a version + tool overrides |
| GET | /api/runs, /api/runs/:runId | run history (?parentRunId=, ?limit=, ?offset=) |
| GET / POST | /api/tools | tool definitions |
| PUT / DELETE | /api/tools/:id | update / soft delete |
| GET | /api/tools/implementations | registered code implementations |
| POST | /api/prompt, /api/prompt/stream | ad-hoc run with a tools list |
| GET | /api/models | effective tier → model mapping |
| PUT / DELETE | /api/models/:tier | repoint / drop a tier binding |
| GET | /api/brief | this install's authoring brief, generated from the live registries |
| GET | /api/files?path= | serve generated files |
| GET | /api/health | health check |
Requests validate against zod schemas in src/domain; failures return 400 with
{ error, fields } keyed by dotted path. That's what keeps a malformed graph
from persisting and only exploding later inside the runner.
Tools
Built in, thirty-three of them:
| | |
|---|---|
| Files | read_file, write_file, list_files, search_files |
| Code | run_script — model-written Node/Python/sh in a sandboxed subprocess |
| Git | git_log, git_diff, git_show, git_status, git_blame, git_ls_files over any repository under a data directory — . for the one the run was started in; git_clone, git_checkout, git_add, git_commit, git_push confined to the run's workspace |
| The world | web_search, http_fetch, ocr, generate_image |
| Run state | state_get, state_set, state_list — dies with the run |
| Durable memory | memory_get, memory_set, memory_list — outlives it |
| People | ask_human, interview, await_external — these park the run |
| Peers | send_message, receive_message, list_peers |
| Self-critique | self_reflect |
Eight of them are not directly assignable — generate_image, web_search,
ask_human, interview, git_commit, and the three memory_* — because they
refuse to build without config bound on a tool definition: a model, an engine, a
question, a brief, an author identity, a namespace. A gate without its wording and a memory
without its namespace are not usable tools, so they only exist as rows.
Writing your own takes no fork and no rebuild. Drop a module in ./tools
(TOOLS_DIR) — every module in that directory is a provider, so a file in the
right place is the registration:
// tools/whatsapp.ts
import { z } from "zod";
import { defineTool } from "@upship/agents/tools";
export default [
defineTool({
name: "parse_whatsapp",
description: "Parse a WhatsApp export into messages. Pass `after` for only what is newer.",
parameters: z.object({
path: z.string().describe("Path to the export, relative to the data directory."),
after: z.string().optional().describe("Cursor from a previous call."),
}),
async execute({ path, after }, ctx) {
return parse(await readFile(await ctx.files.read(path), "utf8"), after);
},
}),
];Hand defineTool a zod schema and it converts to JSON Schema for the model
and keeps the parser, so a graph node calling your tool gets its args coerced
and validated the way a built-in's are. createToolContext() ships beside it,
to call the tool in a test the way the platform will. The full page is
/writing-tools in the UI, and there is a section in agent brief.
The other two routes, for tools that already exist somewhere:
TOOL_PROVIDERS=@acme/agent-tools,./my-tools.js # packages, which have no directory to sit in
MCP_SERVERS='[{"name":"fs","command":"npx","args":["-y","@modelcontextprotocol/server-filesystem","/tmp"]}]'Provider tools land in the same registry as the built-ins, so the three-tier name
lookup is unchanged. A provider cannot shadow an existing name —
registerSpec() refuses and records the skip — and one that fails to load is
skipped rather than fatal. What loaded, what was skipped and what threw is on
the Tools page, in agent tools, and at GET /api/tools/providers: a provider
that failed and a provider you forgot to write look identical otherwise.
Adding a first-party tool — what we do, because we ship the file — is three
steps: write the factory in src/tools/<domain>-tools.ts, register it in
src/tools/index.ts with a name and description, and it's assignable to any
agent. From outside this repo that route means a fork, and a fork means every
upgrade is a merge.
Tool definitions are separate rows pointing at an implementation, carrying their own description and config. Two definitions over one implementation is how you A/B a tool description or bind a runtime parameter without duplicating code.
Storage
Everything persists through the AgentStore port in src/store/types.ts — a
repository over four aggregates, not a query builder. Pick an adapter at boot:
- postgres (default) — drizzle/pg; tables are created on first start
- sqlite — Node 24+,
node:sqlite, zero extra dependencies;SQLITE_PATH - filesystem — one JSON file per record under
FILE_STORE_ROOT, atomic writes - memory — what the test suite runs on
Two rules make the non-SQL adapters possible: application code generates ids,
timestamps, and version bumps before the store call, and a missing record
returns null rather than throwing. The domain model is owned by zod in
src/domain, deliberately not derived from the Postgres schema — that's the
coupling the port exists to break.
Adding an adapter means implementing the port and passing
src/store/__tests__/conformance.ts. Nothing else.
Embedding the runtime
The third way to run this, after a host and a laptop: as a dependency of your
own application, where runAgent is a function your code calls, the store is
your database, and there is no HTTP between you and it. From your project:
agent bundle # writes lib/agents from the install you are running
npm install ./lib/agentsimport { createRuntime, createAgent, runAgent } from "@upship/agent-core";
const rt = await createRuntime({ driver: "sqlite" }); // store, tools, models, liveness timers
const agent = await createAgent({ name: "hello", description: "says hi", systemPrompt: "Be brief.", tools: [] });
const run = await runAgent(agent, "hi", rt.ctx, rt.models);
await rt.stop();createRuntime is the same boot the API server does, so an embedded runtime is
the server's runtime minus the HTTP. The bundle carries both the TypeScript and
the compiled output, so nothing compiles on install; re-running agent bundle
after upgrading the CLI is the update path. The generated lib/agents/README.md
covers the options and the environment variables.
The worked example — register a tool, create an agent, run it, read the run
back, with a paragraph on what each line commits you to — is the Embedding
page of the web UI (/embed, beside Graphs), and the script it walks through is
docs/embed/example.ts, which runs as written against
a fresh bundle. That page also says which path to take (agent bundle for
source you read and patch; npm install @upship/agents once the package is on
the registry, which it is not yet), what an embedder owns, what it doesn't need
(auth, projects, the thin client are the service's problems), and why the UI is
not embeddable.
Versioning and runs
Every save snapshots a version with an auto-incremented patch. Runs record the
agent version that produced them, so any run can be replayed against a different
version or with tools swapped. Run history is a collapsible tree — sub-agent
calls hang off their parent via parentRunId — with token, time, and cost
rollups across the whole pipeline.
Replay is late-bound on purpose: an old version runs on whatever its tier points
at now, matching how it already picks up current tool definitions. Pin by
setting model on the agent.
Testing
npm test # full suite, memory adapter, no database needed
npm run test:watch
npm run test:adapters # application suite once per adapter
npm run test:adapters sqlite # or just oneThe Postgres adapter test talks to pg directly and skips with a warning when
none is reachable — opt-in on your machine, required in CI. .github/workflows/ci.yml
typechecks and builds, then runs the suite on Node 24 and 26 against a Postgres
17 service, failing rather than skipping if pg was unreachable.
Layout
src/
domain/ zod schemas — the model, owned here, not by any storage schema
api/ Hono server, routes, request validation
runner/ execution: freeform loop, graph DAG, reflection, structured output
registry/ agent + tool registries, storage-agnostic application logic
store/ the AgentStore port and its four adapters
tools/ tool implementations, the provider port, MCP client
cli/ the `agent` command
ui/ React + React Router + Tailwinddocs/agentic-design-patterns.md covers the patterns behind the execution modes.
CLAUDE.md is the working design record — an index over docs/arch/, which is longer and closer to the code.
