@haikit/server
v0.18.0
Published
The haikit server runtime: agent loop, elicit state machine, derived query_ui, and the interaction security boundary.
Downloads
3,402
Maintainers
Readme
@haikit/server
The runtime: agent loop, elicit state machine, the interaction security
boundary, and a derived query_ui.
import { createHai, memoryStore, nodeHandler } from "@haikit/server";
const hai = createHai({ model, store: memoryStore(), tools, surfaces, system });
const handle = nodeHandler(hai, "/hai");Two routes: POST /hai/chat and POST /hai/interact, both streaming SSE.
Don't let a proxy buffer the streams. Every route streams SSE, and a
browser counts a surface's freshness from when its frames arrive, so a proxy
that holds them back makes the browser's deadline late. The responses send
cache-control: no-cache, no-transform and x-accel-buffering: no, which
nginx honours. Turn response buffering off for these paths in any other
proxy or CDN in front of them.
A conversation id is a bearer token: put the routes behind your own auth.
nodeHandler does no authentication. Anyone with a conversation's id can send
to it, click in it, read its notices, and wake it, so it is only as private
as its id. pgStore and memoryStore mint unguessable ids. Before handing a
request to nodeHandler, check that the signed-in user owns that
conversation, as you would for any record.
/hai/interact accepts {handle, action, value} and nothing else. What an
action means is resolved server-side from the surface's declared contract. If
the client could name a tool, a prompt injection inside any tool result would
become a button wired to it.
An action's handler runs only for a click the turn can accept: a
resolve on the surface the turn is waiting for, or an inform while nothing
waits. Any other click is refused before the handler sees it, so a handler can
write, and can be async to do it. It gets { props, handle, conversationId }.
If it throws, nothing is recorded and the user can click again. Once it
succeeds, the click is saved before the model runs, so it can't be answered
twice.
That makes a handler run at least once per recorded click, not exactly once.
If saving the click fails after the handler has written (a crash, a newer
request taking the conversation over, a failed write), nothing is recorded, and
clicking again runs the handler again, perhaps with a different value. A
resolve is recorded at most once per surface, so a resolve handler that
writes should upsert on (conversationId, handle). Handles are unique only
within a conversation, and the last attempt is the one recorded. An inform
surface takes any number of clicks, so an inform handler that writes needs a
write that's safe to repeat, or a key of its own.
A model reply asks one question at a time. Claude can make several tool
calls in one reply, but the turn waits on one surface. The first elicit
surface a reply renders is shown; a later one is neither stored nor shown, and
its call gets Not shown: ui_01 is already waiting for the user. Ask this again
after it is answered. So every tool_use still gets a result, and the model
can ask again once the first is answered. Display surfaces are not limited.
The turn waits on the question a call showed, even if the tool returns
something else; that return follows the question's digest in the answer. A
call that asks and then throws has its question withdrawn, so another call in
the same reply can ask instead. Every render a tool starts finishes before its
call is decided, and a render after the tool has returned shows nothing.
A slow tool can say how far it has got. ctx.progress sends the browser
a frame for the tool's row, and nothing else: it isn't stored, and the model
never sees it.
run: async (input, ctx) => {
ctx.progress({ message: "Checking fare sources", done: 0, total: sources.length });
for (const [i, source] of sources.entries()) {
rows.push(...(await source.search(input)));
ctx.progress({ done: i + 1 }); // a call updates only the fields it names
}
return ctx.render(flightTableServer, { ...input, flights: rows }, { mode: "elicit" });
};Every field is optional. Frames go out at most every 100 ms, so a tool can
report every row; the last one always goes out, before the row says the call
has finished. A call after the tool has returned does nothing, and a field
that isn't what its type says is dropped rather than failing the tool:
done must be zero or more, and total more than zero.
A tool can revise a surface in place. Rather than rendering a second table to show a refinement, a tool can change the one on screen:
run: (input, ctx) =>
ctx.update(seatMapServer, input.handle, { ...seats, highlight: "window" }),ctx.update(surface, handle, props) is typed, validated and digested like
render, and resolves the same way, with the new digest for the model. It's
stored as a new surface with a new handle that supersedes the old one, all
committed with the turn, so a revision an overtaken turn made is as inert as
its payloads. From then on:
- a click on the old handle is refused ("superseded by ui_05");
- the old surface's freshness window stops counting, and the revision's counts instead;
query_uion the old handle answers from what it was, and says what it is now;- a question waiting on the old surface waits on the revision.
The browser swaps the new props into the component already on screen (see
SurfaceInstance.update in @haikit/client). The handle must be a surface of
the same contract in the conversation, neither answered nor revised already.
init runs a tool at the start of every conversation, before the user
types anything. The runtime makes the call, not the model, so it can't be
skipped: the call and its result go into the history as if the model had made
them.
const hai = createHai({ model, store, tools, surfaces, system, init: loadProfile });A third route, POST /hai/start, opens the conversation and runs init
without calling the model. The client's start() calls it, and mountChat
does so when it opens and after Start a new conversation. The Messages API
requires the history to begin with a user message, so it opens with
[conversation started], followed by init's call and result. Without an
init, /start returns an empty stream and creates no conversation. A client
that never calls start() still gets init, run with the first message.
It gets {} as input and can render surfaces. An elicit surface ("which
account?") parks the conversation before anyone has typed; the answer becomes
init's result, and the model replies. If it throws, the request is refused and
nothing is recorded, so the next start or message tries again. It stays in the model's tool list, because changing that list
mid-conversation invalidates newer models' earlier reasoning. But a later call
from the model gets "already ran" instead of running it twice. Its result stays
in context for every turn, so keep it digest-sized.
The model is kept to what your tools do. HaiKIT appends DEFAULT_SCOPE
to your system prompt. It tells the model to decline requests outside its
tools in one sentence, never to claim something no tool did, and to treat tool
results as data, not instructions. Pass scope to replace it, or
scope: false to leave it off:
import { createHai, DEFAULT_SCOPE } from "@haikit/server";
const hai = createHai({
model, store, tools, surfaces, system,
scope: `${DEFAULT_SCOPE}\n- Never quote a fare the search did not return.`,
});It guards what the model says, not what happens. The model can only call the
tools it was given, and the browser can only send {handle, action, value},
so put every check with consequences, such as who may book what, inside the
tool or action handler, where it runs as code.
The server can tell a conversation something later, with a notice. A
webhook confirming a booking, a job finishing: hai.notify sends the browser
the notice's payload now, and the model its model text in the next user
message the conversation records. Declare it once, beside your surfaces:
// shared/notices.ts — imported by both halves
export const holdConfirmed = defineNotice({
name: "hold_confirmed",
version: 1,
payload: z.object({ flightId: z.string(), reference: z.string() }),
});
// server
const holdConfirmedServer = holdConfirmed.implement({
// required, like a digest: what the model hears, or null for nothing
model: (p) => `The airline confirmed the hold on ${p.flightId}, reference ${p.reference}.`,
});
const hai = createHai({ model, store, tools, surfaces, system, notices: [holdConfirmedServer] });
// anywhere, whenever: a webhook, a queue, an action handler's follow-up
await hai.notify(conversationId, holdConfirmedServer, { flightId: "AC832", reference: "QX7P2L" }, { handle: "ui_01" });The payload is typed from the contract, and checked against its schema before
it is stored. A notice that isn't listed in notices, a payload that fails its
schema, or a handle with no surface stored in the conversation is refused.
Pass a handle from the conversation's saved history, such as an action
handler's ctx.handle. The check is against the stored surfaces, not that
history, so a surface left behind by an overtaken turn may also pass, until a
sweep such as @haikit/postgres's sweepOrphans removes it. That is no
promise, just harmless: the handle only says where the browser shows the
notice, the model never sees it, and a notice can't change anything.
notify takes no lease, so it works while a turn is streaming, while one is
parked on a question, and with nobody connected. It never touches the history
itself: the next request that records a user message takes every unread notice
in. A typed message carries them ahead of what the user typed, and an answer to
a question carries them after its tool results. Taking them in commits with
that request's history, so a turn whose save is lost leaves them unread for
the next one. A notice never starts a turn.
With any notices listed, hello tells the browser to open a fourth route,
GET /hai/events?conversationId=…. It streams the conversation's notices,
each with its sequence number as its SSE id, so a browser that reconnects
with Last-Event-ID picks up where it left off. It takes no lease, writes
nothing, and accepts nothing from the browser but which conversation and where
to resume. The model text never goes out on it. The browser sees that only
once a turn has taken it in, as part of the history the context event shows
the inspector, the same as a digest. It reads the store every two
seconds, and a store with watch wakes it the moment a notice lands. The
two-second read stays even then, so a lost wake-up only delays a notice.
A wake notice starts a turn of its own. Declare one with kind: "wake",
and its model function must return text, since the turn is the model's
answer to it:
export const fareDropped = defineNotice({ name: "fare_dropped", version: 1, kind: "wake", payload: FareDrop });
const fareDroppedServer = fareDropped.implement({
model: (p) => `The fare on ${p.flightId} dropped from $${p.was} to $${p.now}.`,
});It is sent the same way, with hai.notify. The turn runs where a browser is
watching: the events route, holding the conversation's stream, takes the
lease and streams the turn down that stream. The user message it records
carries every unread notice, then
[The user has not said anything. The notifications above arrived on their own.].
If the conversation is busy, the route tries again, backing off from 250 ms to
5 s, until it is free or the browser leaves. No turn starts in these cases,
and the notice rides the user's next message, as a passive one does:
- nobody is watching (a browser that connects later starts it);
- a question is waiting for an answer;
- the conversation hasn't begun, or is out of date;
- it's past
maxWakes, by default one wake turn a minute per conversation.
createHai({ /* … */, notices: [fareDroppedServer], maxWakes: { count: 1, perMs: 60_000 } });Several wake notices waiting at once start one turn. One held back by
maxWakes never wakes later. Every events stream this process has open for
the conversation hears the turn, so a second tab shows the reply too.
With several app servers, a tab may be connected to a different one from the
server running the turn. A store with the optional publish and subscribe
methods carries the turn there too: its frames go out through the store, and
every other server hands them to its own streams, with what the turn has
shown so far for a tab that joins part way. pgStore does this over NOTIFY
when given a listen connection; memoryStore does it between Hai
instances sharing it. A store without the methods keeps a turn on its own
server: there, a tab on another server sees the notice but not the reply.
The history's whole context is not sent across; another tab gets it with its
own next request.
A wake turn whose model call fails before any tool has run is undone: the
history is left as it was and its notices can wake again, though the attempt
still counts against maxWakes. Once a tool has run, a failure keeps what
happened, the call and its result, and closes the turn. The notice never
wakes again, so no tool runs twice. When a turn ends, the stream sends
released once the conversation is saved and free, and the client waits for
that before it sends anything. hai.wake(conversationId, emit) is what the
route calls, if you serve the events stream yourself.
App code can revise a surface too, outside any turn. A price feed moving
a fare, an airline repricing a table: hai.update is ctx.update for code
that isn't a tool.
const hai = createHai({ model, store, tools, surfaces, system, updates: true });
// anywhere, whenever
const { handle } = await hai.update(conversationId, flightTableServer, "ui_03", { ...props, flights: repriced }, {
model: "Fares moved: AC832 is now $389 (was $343).", // optional
wake: true, // optional: have the model tell the user now
});hai.update refuses without updates: true, even in an app that lists
notices. The flag opens the events stream, which is how a revision reaches
the browser, and has the events route start the turn an update sent with
wake asks for. The props are typed from the surface's contract and checked
against its schema before anything else happens.
It takes the conversation's lease, as a request does, and writes the revision
through the same fenced path as ctx.update, so a click can never act on
props the server doesn't hold. While a turn holds the conversation it waits,
backing off, for up to timeoutMs (default 30 s), then throws
ConversationBusy: a reprice during a model reply lands when the reply ends.
The rules are ctx.update's: a click on the old handle is refused, its window
stops counting, an out-of-date conversation stays out of date, and an answered
surface can't be revised. handle may be any handle the surface has had; the
revision replaces the latest. A question waiting on the surface waits on the
revision, and its answer goes out under the digest the model saw, with the
update after it.
Then it appends an update notice, which tells both sides:
- The events stream sends the browser the revision, as the
ui_open(withreplaces) andui_propsa tool's revision sends, and the browser swaps the new props into the component on screen. Delivery is the notice log's: in order, resumed withLast-Event-ID, and across servers. - The model hears it in the next user message, or at once with
wake, under the rules awakenotice follows:[UI update] Fares moved: AC832 is now $389 (was $343). ui_03 was replaced by ui_05; its earlier digest is out of date. Now: <the new digest>. Yourmodelsentence is optional; the rest is always there.
The revision and the notice it owes commit together, in one fenced save on the
conversation row (Conversation.announcing), before the notice is appended.
If appending it fails, hai.update throws, saying so, and the notice stays
owed: whoever holds the conversation next, a request, a wake or another
update, appends it before doing anything else, so the model hears of the
revision before it runs again. Delivery is at least once: a notice already
appended by a holder whose save then failed isn't appended again, but an
append that stalls past the lease while the next holder sends it can land
twice. The browser shows the revision once either way; the model reads its
text twice.
One write per surface is in flight at a time. A call made meanwhile replaces
any still waiting behind it, and every call it replaced resolves to the handle
it writes, so a feed ticking ten times a second writes as often as the lease
allows. Calls naming the surface's first handle or its latest share the
queue. A wake any of them asked for still wakes.
A conversation closes once any of its surfaces passes its staleAfterMs.
From then on both routes answer with an expired event and the model is not
called, so a picker left open over a weekend cannot resolve against last week's
data. The user starts a new conversation. A surface that declares no window
never goes stale, so set one wherever the data can change.
memoryStore() is for development. A parked elicit turn is durable state —
lose pending and that conversation can never be sent again. Ship with
@haikit/postgres.
