@jsm-mit/chat-agent-package
v0.7.0
Published
Agent framework: a configurable LLM-driven Agent class plus message-bus agent types (personal assistant, summarist, marketing), with a persistent CLI sandbox game for testing personas.
Readme
chat-agent-package
An agent framework in two layers.
The core is the configurable, instantiable Agent class
(personality/system prompt, conversation memory, tool calling,
swappable model provider — OpenAI current-gen for now), plus a
persistent CLI sandbox game for testing agent personas interactively.
The agent types on top of it are the ones a host wires to real
traffic: PersonalAssistantAgent, SummaristAgent, MarketingAgent,
WebChatAgent.
They live here rather than in the service that runs them, so one type's
behaviour can be read — and tested — on its own.
Part of workspace-agent-platform; concept/architecture design lives
in ~/workspace-jsm-company/jsm-ideas/agent-platform/.
Status
Agent class + OpenAiProvider + CLI sandbox game + the message-bus
agent types, all with tests.
Agent types and the bus
A host merges every source it listens to into one stream of
InboundEvents — a message, a heartbeat tick, or an assigned task —
and agents subscribe to it themselves. There is no router deciding who
gets what:
sources ──┐
├──► inbound$ ──► agents subscribe and filter themselves
heartbeat ┘ │
▼
outbound$ ──► the host's transportsChannelAgent is the base every type extends. It owns the
subscription, matching an event to the agent's sources, the
claim-or-share call, the error isolation and the staleness split, so a
subclass is a couple of async handlers (plus an optional synchronous
accepts() filter) — no rxjs in sight, and a unit test drives it with
plain objects.
Every agent is built with an AgentBinding: the sources it hears (a
source name plus one endpoint, an endpoint prefix, or a task name) and
whether it hears the heartbeat. Each source says whether the agent
claims what arrives there or shares it, and which channel an answer
goes out on — absent, the agent stays silent there.
Exclusivity without a central switch: the bus delivers each event to every subscriber synchronously, in subscription order. The first agent that claims an event takes it, and nobody started later sees it; an agent that shares handles the event and lets it pass on. Priority is therefore the order the host starts its agents in — one ordered list, in one file.
ConversationalAgent adds an Agent per conversation with its history
restored and written back through a store. A type that watches rather
than converses — SummaristAgent — extends ChannelAgent directly and
calls the provider for a one-off summary, so an LLM-less agent is an
ordinary subclass rather than a special case.
Everything naming a transport or an endpoint is an opaque string here. This package still knows nothing about WhatsApp: how an agent answers comes from its own binding and settings, and the host turns outbound events into real sends.
Setup
npm install
cp .env.example .env # then fill in OPENAI_API_KEY
npm testSandbox
npm run sandboxAn arrow-key menu (state lives in sandbox/db.sqlite, gitignored):
create an agent (name, persona, whether it discloses being an AI),
pick a saved agent to chat with, edit one, or delete one. Chatting
resumes the agent's prior conversation automatically; /exit (or
Ctrl+D) ends the session and saves it.
Editing picks one field at a time (name, persona, disclosure, tools). Name and persona prefill the input line with the current text so you edit it in place — backspace/insert — instead of retyping it whole.
Tools
An agent can be given Tools (name/description/JSON-schema
parameters/execute, plus an optional progress line a chat may show
while the tool runs) it calls mid-conversation via real OpenAI
function calling — e.g. "what's your availability tomorrow?" triggers
a check_availability call, the result feeds back into the model,
which then replies naming an actual slot. execute is a plain async
function — nothing stops it from calling a canister through its
wrapper package instead of returning canned data.
The sandbox ships one fake tool, check_availability
(sandbox/tools.ts) with no real backend — just to exercise the
mechanism. Assign tools to an agent via the "Tools" field in create/edit
(checkbox-style toggle menu). AgentConfig.toolNames is declarative
(persisted in SQLite) and resolved against a tool registry at Agent
construction time — the registry, not the config, decides what
execute actually does, so the same config can point at fake data in
the sandbox and a real backend in production.
WebChatAgent — the chat is the memory
WebChatAgent is a member of a multi-party chat that lives in a transport of
its own (a chat canister). It extends ChannelAgent directly, not
ConversationalAgent: it keeps no history. Every turn it reads the chat back
through the ChatTranscriptSource port, hands the model everything up to the
lines it has not answered yet, and answers those — so what humans said while
the agent was muted (agentsReply: false) is in its context the moment it
speaks again, and a restart loses nothing.
Inbound: from is the chat id, senderId the author, messageId the chat's
sequence number; turns on one chat run in order. Outbound: send-text with
no name prefix (the chat shows the author), and send-trace lines for
transports that keep an audit trail next to the conversation: a tool's own
progress line ("Sprawdzam listę usług…") before the tool runs, so the chat
shows what the agent is doing while the person waits, and — when the host
supplies traceOf — a line per finished tool call. Traces never reach the
model: the transcript's tool traces are left out of its context.
Remembered tool exchanges
The transcript holds what was written, not the tool calls behind a reply, so
a tool's result would be gone one turn later. That breaks a two-step write:
turn N calls a tool with confirmed: false and gets a preview with a
confirmationId; turn N+1 the owner says "tak", and the model must call again
with that id — which it can only do while it still sees turn N's exchange.
One agent, many chats, a prompt of its own for each. personaFor(conversationId) on the
context is asked at every turn and its answer is that turn's system prompt; undefined, or no
hook at all, gives the agent's persona. It mirrors ToolProvider.toolsFor: a host whose agent
serves every customer in a chat of their own can tell it, per chat, which language to answer in.
So WebChatAgent keeps, per chat, the tool calls and results of its last
MAX_REMEMBERED_TOOL_TURNS (3) turns that called a tool, each pinned to the
newest line its turn answered (its transcript seq). A turn without tool calls
keeps nothing. On the next turns the history is built from the transcript and
cut to historyLimit first; then each exchange goes back in whole, right before
the first line newer than its anchor — normally the agent's own reply:
owner: dodaj strzyżenie za 80 zł ← anchor
assistant: (calls add_service) ┐ remembered
tool: {confirmation_required, …} ┘
assistant: Dodam usługę… Potwierdzasz?
owner: tak ← this turnAn exchange whose reply is not in the chat yet goes last, unless its line is
still pending; one whose line fell out of the historyLimit window is left
out. A tool result therefore always directly follows the call that asked for
it, as the OpenAI provider requires. Exchanges come on top of historyLimit.
The memory is process-local, bounded per chat and to MAX_REMEMBERED_CHATS
(1000) chats, the chat written longest ago dropped first. A restart forgets it,
which is acceptable: without the old confirmationId the model asks the write
tool for a fresh preview, and the owner confirms once more.
