@chatcore/server-sdk
v0.1.0
Published
TypeScript server-side SDK for the chat backend v1 API — runs on your backend, where the project api_secret lives.
Readme
@chatcore/server-sdk
TypeScript server-side SDK for the chat backend. It is meant to run on your
own backend, where the project's api_secret lives — never in a browser or
a mobile app.
Licensed Apache-2.0 — the full text ships inside the package (LICENSE), so
it travels with every install. Everything below is written for somebody who
installed this package from npm and cannot see the repository it is built in:
the compatibility table and the SemVer rules are what each release is held to.
The release number itself lives in package.json and is deliberately not
repeated here, because a version spelled out in prose goes stale at the next one.
Project and tenant are two different things
M14 (spec 20 §26.1) split what used to be one word:
- A project owns the credential —
api_key,api_secret, rotation, and the tier-1 disable switch. It is what the platform operator provisions, and it is the only place a credential exists at all (§26.2). A tenant has none. - A tenant is the isolation boundary. Every user, channel, message, upload and event hangs off exactly one tenant, and two tenants of the SAME project are as isolated from each other as two customers (§26.5) — "same customer, so let them see each other" is not a rule this API has. A project holds up to 100 tenants, disabled ones included.
So tenantId is required in ChatClientOptions and is a tenant UUID, not a
slug and not the project id. The SDK stamps it as the tid claim on both mint
paths, and the server refuses any token without one. Passing something that is
not a UUID raises ChatConfigError at construction rather than a blind 401 on
the first request.
Where the first one comes from: provisioning a project always creates a
default tenant in the same transaction, and POST /v1/admin/projects answers
{project, apiSecret, defaultTenant, created} — so the operator who hands you
apiKey and apiSecret hands you defaultTenant.id alongside them, and that is
your tenantId. It is not a convenience: tid is mandatory on every token, so a
project with no tenant could not sign a request, including the request that would
have created its first tenant. After that, chat.tenants.create makes more of
them and forTenant acts as them.
Quickstart
import { ChatClient } from '@chatcore/server-sdk';
// The custom-data shape of THIS tenant. Every resource is generic over it, so
// `user.custom.department` is typed rather than `unknown`.
interface AppData {
user: { department?: string };
channel: { ticketId?: string };
}
const chat = new ChatClient<AppData>({
apiKey: process.env.CHAT_API_KEY!, // the PROJECT's key
apiSecret: process.env.CHAT_API_SECRET!, // never sent; signs tokens locally
tenantId: process.env.CHAT_TENANT_ID!, // REQUIRED — tenant UUID, not a slug
baseUrl: 'https://chat.example.com', // origin only — the SDK adds /v1
});
// A user token for one of your users. SYNCHRONOUS: HS256 over node:crypto, no
// request, no await. Default lifetime 24h.
const token = chat.createToken('u1', { expiresIn: '2h' });
await chat.upsertUsers([{ id: 'u1', name: 'An', custom: { department: 'design' } }]);
const channel = chat.channel('group', 'design-team');
await channel.getOrCreate({ name: 'Design', members: ['u1', 'u2'], createdBy: 'u1' });
const message = await channel.sendMessage({
userId: 'u1',
text: 'Ready for review',
textFormat: 'markdown',
});
// Messages page by sequence, not by cursor.
const page = await channel.queryMessages({ limit: 20 });
for (const m of page) if (!('deleted' in m)) console.log(m.seq, m.text);
const older = await channel.queryMessages({ limit: 20, beforeSeq: page.at(-1)?.seq });
// The query read models page by cursor.
const users = await chat.queryUsers({ filter: { active: true }, limit: 50 });
if (users.nextCursor) await chat.queryUsers({ limit: 50, cursor: users.nextCursor });Four things about that snippet that are load-bearing:
tenantIdis not optional and not a slug. It is the tenant UUID this client acts as, stamped astidinto every token it mints — bothcreateTokenand the per-request server token. There is deliberately no "the project has one tenant, so infer it" branch anywhere: not in the SDK, not in the server's verify, not on the guest path (§26.3/§26.8).baseUrlis an origin, never a path. The/v1prefix and every URL template come from the operation registry in@chatcore/contracts; no façade method formats a URL (Invariant 1).- A list method returns
Page<T>, which IS anArray<T>with an extranextCursor. Destructuring andfor…ofwork;.map()gives a plain array back, so the cursor cannot be lost by accident — read it off the page. - Everything public is camelCase. The wire is snake_case and exactly one
converter crosses the line (
src/internal/case.ts). Yourcustomkeys, and the other seven skip-set subtrees, cross byte-for-byte.
A server token is minted per request (5-minute TTL, cached, re-minted under 60s
of remaining life) and sent as Authorization: Bearer …. You never construct
one. That cache is keyed by (apiKey, tenantId) — see forTenant below for
why a flat per-key cache would be a security bug rather than an optimisation.
forTenant — one backend, many tenants
const support = chat.forTenant('7c1b…-…-…'); // another tenant of the SAME project
await support.channel('group', 'triage').sendMessage({ userId: 'u1', text: 'hi' });forTenant(tenantId) inherits everything else — apiKey, apiSecret,
baseUrl, timeoutMs, masterApiKey, validateResponses, fetchImpl — and
only tid differs. (There is nothing to inherit for retries: no client-level
retry setting exists. The registry's retry column decides per operation, and
RequestOptions.retry overrides it per call.) The same UUID validation applies,
so a slug is a ChatConfigError here too.
The derived client does not share minted tokens with the one it came from.
The server token cache is per client instance, which makes its effective key
(apiKey, tenantId) — spec 20 §26.9 requires exactly that, because a cache keyed
by apiKey alone would hand a request for workspace A a token that says
tid = B, and every write would land in the wrong tenant while looking perfectly
healthy. test/token.test.ts pins it at the unit level and
apps/api/test/int/sdk/sdk-tenants.int.test.ts pins it against a live server by
interleaving calls from two clients and asserting which tenant each resource
landed in.
There is no cross-tenant read anywhere. "Switch tenant" means "mint a token with
a different tid", which is all forTenant is. The single exception in the whole
SDK is chat.tenants.list(), which answers tenant metadata and never chat
data.
chat.tenants.* — the customer's own control plane
Tenant management is authorised by the project's server token, not by the
platform masterApiKey: holding the project's api_secret is what proves the
right to manage the tenants inside it. The client's own tenantId says which
tenant it acts as on data routes; on these three it only identifies the project,
so chat.tenants.patch(someOtherTenantId, …) is the normal case.
create({slug, name, settings?})→{tenant, created}. Get-or-create by slug within the project: a replay with the same payload answerscreated: false, a replay whosename/settingsdisagree with the live row is409 idempotency_conflict, and the 101st tenant is422 tenant_limit_exceeded.list()→Page<Tenant>, sorted by slug, disabled ones included. No cursor — the cap of 100 is what makes that sound.patch(tenantId, {name?, settings?, disabled?, fcmServiceAccount?}). This is where per-tenant Firebase credentials and guest settings are configured.fcmServiceAccount: nullclears the credential; no route ever echoes it back.
Two disable tiers, and they are independent. chat.admin.patchProject(id,
{disabled: true}) is tier 1 and stops every tenant of the project;
chat.tenants.patch(id, {disabled: true}) is tier 2 and stops one. Re-enabling
the project does not clear any tenant's own flag — a tenant switched off in tier 2
stays off until tier 2 turns it back on.
disabled: true on the last active tenant of a project is 403 forbidden,
transactionally, so a project can never reach zero active tenants and become
unable to mint a working token. Two crossing disables cannot both succeed. A
single-tenant project therefore cannot switch itself off through self-service;
that needs patchProject({disabled: true}) from the operator. Disabling the
tenant this very client acts as is allowed while another is active — it kills
this client's own token, so continue on chat.forTenant(<a tenant still active>).
Disable does not force-disconnect realtime. Both tiers block REST and new connect-tokens immediately, but a Centrifugo token already issued keeps its WebSocket receiving publications until it expires — a tail of up to 1 hour (§26.4). Treat disable as "no new work", not as "the socket is closed now".
Guest creation, and the honest bound on its rate limit
chat.createGuest is the SDK's server-token path: 100 guests per hour per
tenant, unchanged by M14, because a guest quota is per isolation boundary.
The other path is the browser widget's, which the SDK does not use: an
unauthenticated POST /v1/guests carrying the project's semi-public api_key
and now a mandatory tenant_id. That one is limited to 10 per hour per
(project, ip), counted before the tenant lookup so that spraying random
tenant_id values costs quota rather than being free (§26.8). Two consequences
worth knowing before you rely on it:
- Two widgets for two tenants of the same project, behind one IP, share one window of 10.
- The limiter is an in-process
Map, so the real ceiling is 10 × number of running instances per hour per IP. A centralised limiter is M15 work.
createToken for a user that does not exist yet
createToken signs happily for any id: it is local HMAC and talks to nobody.
The token is only rejected later, on that user's first request, where
TokenService calls lifecycle.assertActive(tenantId, payload.sub)
(apps/api/src/modules/auth/token.service.ts) and answers 401. That check
is a live database read, so the same token starts working the moment the user is
upserted, and stops working if the user is deactivated — nothing is baked into
the token. Upsert first, then hand out the token.
What each method calls
Every façade method routes through one registry operation, and every operation
classified facade or namespace has exactly one method — 78 of the 88, checked
in both directions by test/export-surface.test.ts, test/layering.test.ts and
the namespace/channel accounts. The other ten are unreachable on purpose: seven
client-only (PATCH /me, POST /connect-token, GET /sync,
GET /sync/checkpoint, GET /search/messages, GET /me/live-locations,
POST /channels/{type}/{id}/leave) and three internal (GET /health, the two
Centrifugo proxy routes). Asking the transport for one of them is a compile
error, and a ChatConfigError at runtime.
ChatClient
| Method | Route |
| --- | --- |
| createToken(userId, {expiresIn}) | (local HS256, no request) |
| createGuest(input) | POST /guests |
| upsertUsers(users) | PUT /users |
| partialUpdateUsers(updates) | PATCH /users |
| queryUsers(input) | POST /users/query |
| getUser(id) | GET /users/{id} |
| deactivateUser(id) · reactivateUser(id) | POST /users/{id}/deactivate · /reactivate |
| queryChannels(input) | POST /channels/query |
| getMessage(id) · updateMessage(id, patch) · deleteMessage(id) | GET · PATCH · DELETE /messages/{id} |
| channel(type, id) · directChannel(a, b) | (lazy handle, no request) |
Channel (and DirectChannel, which extends it)
| Method | Route |
| --- | --- |
| getOrCreate(input) | POST /channels |
| get() · update(patch) · delete() | GET · PATCH · DELETE /channels/{type}/{id} |
| freeze() · unfreeze() | POST …/freeze · /unfreeze |
| hide(input) · show() · archive() · unarchive() · pin() · unpin() · mute(input) · unmute() | POST …/hide · /show · /archive · /unarchive · /pin · /unpin · /mute · /unmute |
| markRead(input) · markDelivered(input) | POST …/read · /delivered |
| queryMembers(input) | POST …/members/query |
| addMembers(input) · removeMembers(input) | POST …/members · …/members/remove |
| updateMember(userId, patch) | PATCH …/members/{userId} |
| transferOwner(input) | POST …/transfer-owner |
| sendMessage(input) · sendStaticLocation(input) · startLiveLocationSharing(input) | POST …/messages |
| sendSystemMessage(input) | POST …/system-messages |
| queryMessages(input) | GET …/messages |
| pinMessage(id) · unpinMessage(id) | POST /messages/{id}/pin · /unpin |
| queryPinnedMessages(input) | GET …/pins |
| createPoll(input) | POST …/polls |
| queryActiveLiveLocations(input) | GET …/live-locations |
| streamMessage(input) — DirectChannel only | POST …/message-streams |
streamMessage is on DirectChannel and nowhere else, as a type-level
restatement of spec 10 §15.4: streaming is a two-member direct surface, and a
Channel<'group'> handle simply has no such method to call. The runtime check
still belongs to the server.
Channel.resolveId() is what turns a DirectChannel into a channel id, and it
costs a POST /channels the first time. Methods whose URL does not name the
channel (pinMessage, unpinMessage) skip it.
StreamWriter (from directChannel.streamMessage(...))
| Method | Route |
| --- | --- |
| append(token) | (buffers locally, no request) |
| flush() | PATCH /message-streams/{messageId} |
| complete() | POST /message-streams/{messageId}/complete |
| fail({code, message}) | POST /message-streams/{messageId}/fail |
Every flush sends a full snapshot {revision, text}, never a delta, and
revision is a decimal string starting at 1. Flushes coalesce: a flush while
one is in the air waits and then sends the newest buffer, and a
minFlushIntervalMs floor (default 150ms) keeps a
token-by-token caller inside spec 10 §15.2's 5–20/s budget. Pass
{minFlushIntervalMs: 0} if you are doing your own pacing.
Namespaces
| Namespace | Method | Route |
| --- | --- | --- |
| chat.channelTypes | list() · get(name) · update(name, patch) | GET /channel-types · GET/PATCH /channel-types/{name} |
| chat.devices | register(input) · remove(token) | POST /devices · DELETE /devices/{token} |
| chat.userGroups | create · query · get · update · delete · addMembers · removeMembers | POST /user-groups · POST /user-groups/query · GET/PATCH/DELETE /user-groups/{id} · POST /user-groups/{id}/members · …/members/remove |
| chat.reactions | add(messageId, input) · list(messageId) · remove(messageId, type, input) | POST/GET /messages/{id}/reactions · DELETE /messages/{id}/reactions/{type} |
| chat.polls | get(id) · vote(id, optionId) · removeVote(id, optionId) · close(id) | GET /polls/{id} · POST /polls/{id}/votes · DELETE /polls/{id}/votes/{optionId} · POST /polls/{id}/close |
| chat.uploads | create(input) · complete(id, input) · refreshPutUrl(id, input) · cancel(id, input) | POST /uploads · POST /uploads/complete · POST /uploads/{id}/refresh-put-url · DELETE /uploads/{id} |
| chat.files | getDownloadUrl(uploadId) | GET /files/{id}/download-url |
| chat.locations | update(messageId, input) · stop(messageId) · queryActive(input) | PATCH /messages/{messageId}/location · POST …/location/stop · GET /users/{userId}/live-locations |
| chat.presence | get() | GET /presence |
| chat.tenants | create(input) · list() · patch(tenantId, patch) | POST /tenants · GET /tenants · PATCH /tenants/{id} |
| chat.admin | createProject(input) · rotateSecret(projectId, input) · patchProject(projectId, patch) | POST /admin/projects · POST /admin/projects/{id}/rotate-secret · PATCH /admin/projects/{id} |
chat.uploads.create/complete are singular in the SDK and batch on the wire:
the route takes up to ten tickets, the façade sends a batch of one and unwraps
it.
The two acting-as-user surfaces, and one that is not
A server token has no user of its own, so routes that act on a user's behalf take one explicitly. Three shapes exist and they are not interchangeable:
- In the body, e.g.
sendMessage({userId}),markRead({userId, seq}),uploads.create({userId, …}). Most of the surface. - In the query string:
uploads.refreshPutUrl(id, {userId}),uploads.cancel(id, {userId})andreactions.remove(messageId, type, {userId}). Without it, the server answers400 invalid_request. - Not at all:
chat.locations.update/stopaccept{userId}for signature symmetry with the rest of the surface and never send it. The server authorises against the share's real owner.
userId is optional in the TYPES throughout, because spec 06's examples call
without it, and mandatory in PRACTICE for a server token on every route that
resolves an actor. Read a 400 invalid_request on a first call as a missing
userId before anything else.
Retry policy
Retries are decided by the registry's retry column, never by the call site.
| Class | Ops | Behaviour |
| --- | --- | --- |
| safe | 74 | Retried on a retryable failure, up to 3 attempts by default. |
| unsafe | 13 | Never retried. Ten are client-only/internal and unreachable; the three you can call are guests.create, userGroups.delete and members.transferOwner. |
| never | 1 — admin.rotateSecret | Hard-blocked. retry: {maxAttempts: 5} does not lift it. |
A retry happens only when all of these hold: the class is safe, an attempt
remains, and the failure is a transport/timeout error or a 408, 429, 502,
503, 504.
500 is not retried. No endpoint is declared safe to repeat after the server
already began handling the request and failed inside it. Neither are 4xx
validation, auth, permission, 404 or 409 — repeating those changes nothing.
Backoff is exponential with full jitter: random() * min(300ms * 2^n, 10s).
A Retry-After header (seconds, or an HTTP-date) is honoured as a floor, and
is not capped by the 10s backoff ceiling — the server's number wins.
One ceiling the SDK adds on top of the spec: a Retry-After longer than
MAX_RETRY_AFTER_WAIT_MS = 60s stops the retry instead of sleeping through
it. The ChatApiError you get keeps its retryAfter, so you can schedule the
work yourself. Without it, 503 + Retry-After: 86400 parks one call for 24
hours inside a 3-attempt budget. A true end-to-end deadline
(RequestOptions.deadlineMs) is deliberately not in 0.1.x — timeoutMs is
per attempt.
Idempotency keys are generated once, before the first attempt, and every
retry resends the same bytes: the body is serialised once, outside the loop.
That covers sendMessage/sendSystemMessage/createPoll/streamMessage
(client_msg_id) and uploads.create (id). Pass your own clientMsgId when
you want the key to survive a process restart, not just a socket error.
Per call: { signal, timeoutMs, requestId, retry: false | {maxAttempts} } as the
last argument of every method. retry: false means exactly one attempt.
Errors that are not failures
Three 409 codes mean "this already happened", and each can mean "you did
it, a moment ago, and the answer got lost". The SDK handles one of them for you
and deliberately does not handle the other two.
stream_closedafter your owncomplete()/fail(). The SDK swallows exactly one case and only insideStreamWriter: the error must carryattempts >= 2(so an earlier attempt of this same call could have landed the write), and a confirmingGET /messages/{id}must show the message in the terminal state this call was asking for —completeaccepts onlystate: 'complete',failonly'failed'. Both conditions are needed: 429 and 503 are retried too, and those attempts provably closed nothing, so a stream the timeout sweeper had alreadyfailedwould otherwise report your generation as finished. Everything else throws:attempts === 1is a genuine conflict, and a caller's secondcomplete()throws — both of those are deliberate, not gaps. Residual and unclosable without the server naming who closed the stream: another actor completing it during a 429 retry.location_not_active/location_revision_conflictafter your ownlocations.update/stop(M12 handoff). The SDK does not swallow these — the CAS is over caller-chosenlocation_seqvalues on a GPS ticker, and only the caller knows whether the seq in the conflict is its own retry or another device's. Treat "conflict on the seq I just sent, with the payload I just sent" as applied, re-read withchannel.queryActiveLiveLocations()if you need certainty, and let a lost update go: the next sample is 1–5 seconds away and supersedes it. Going backwards is silently ignored by the server, and gaps are legal, so there is nothing to repair.created: falseis not an error at all.channel.getOrCreate,chat.tenants.createandchat.admin.createProjectare get-or-create and return the flag.
Compatibility
| | |
| --- | --- |
| SDK | 0.1.x |
| API | v1 — every route under the /v1 prefix |
| OpenAPI document | openapi.json inside @chatcore/contracts, info.version 1.0.0, 88 operations — the API version, which moves on its own schedule and is not the package version |
| Contracts | @chatcore/contracts 0.1.x, the only runtime dependency. Lockstep: the two packages are released together on the same version and this one depends on an EXACT contracts version, so a resolver can never pair them across a release. Do not override it. |
| Node | >= 24 — this package's own engines, matching the repo root; .nvmrc pins 24.13.0. Needs global fetch, node:crypto and AbortController. engines is advisory: npm warns rather than refuses unless you set engine-strict. |
| Module format | CommonJS (main + types from dist/), built with tsc. No ESM build — require is the shape a consumer's node_modules gets. |
| License | Apache-2.0; the full text ships in the package |
0.1.x speaks v1 and only v1. A v2 prefix would be a new major of the SDK,
not a minor.
An SDK is expected to be older than the server it talks to. Response parsing
is Zod, which strips unknown keys, so a server that adds a field does not break
an older client — with validateResponses: true as well as without it, and there
is a forward-compatibility test for both modes. The reverse (an SDK newer than
the server) is not supported: a method for a route the server does not serve
gets a 404.
SemVer
The public surface is src/index.ts and nothing else. src/generated/ and
src/internal/ are not exported, are not part of the promise, and change
whenever contracts change.
| Change | Bump |
| --- | --- |
| New method, new namespace, new operation | minor |
| New optional field on an input or a resource | minor — not breaking |
| New optional parameter on an existing method | minor |
| A response field the server started sending | patch (no SDK edit needed) |
| Renaming or removing a field, method, type or namespace | major |
| Making an optional field required | major |
| Narrowing an accepted type, or widening a returned union | major |
| Changing a default (timeoutMs, minFlushIntervalMs, retry counts, the Retry-After ceiling) | major |
| Changing which operations are safe/unsafe/never | major |
The last two rows are the ones that look like details and are not: a default is
the behaviour of every caller who did not think about it, and the retry class of
an operation decides whether a request can be sent twice. MAX_RETRY_AFTER_WAIT_MS
landed in 0.1.0 for exactly this reason — adding it later would turn a call that
used to hang for 24 hours into one that fails at 60 seconds.
Admin operations and the second wall
chat.admin.* sends the platform masterApiKey as x-master-api-key; a client
constructed without one raises ChatConfigError before any request. A wrong or
missing key is 401 unauthorized from MasterKeyGuard.
That is one of two independent layers. Spec 05 §13 also requires the edge
(Caddy) to route /v1/admin/* only from an IP allowlist, or to keep it off the
public listener entirely. The SDK does not participate in that layer and cannot
get around it. Calling chat.admin.* from an address outside the allowlist
fails at the edge, before the request reaches MasterKeyGuard — and since the
edge does not speak this API's error envelope, the SDK surfaces it as a
ChatApiError with the edge's status and code internal, not unauthorized.
Run admin calls from inside the allowlisted network. (The Caddy configuration is
M15's; there is no deploy/Caddyfile in this repo yet, so the wall described
here is the spec's, not something this repo measured.)
The three admin methods are project-level (§26.7): createProject,
rotateSecret(projectId, …), patchProject(projectId, …). The master key no
longer manages tenants at all — the three master-key tenant routes under
/admin/ were deleted outright, with no alias and no compatibility shim. A grep
gate keeps that path prefix from coming back anywhere in the repo. Creating or
patching a tenant is chat.tenants.* under the customer's own server token, and
rotation moved up to the project.
One consequence, accepted deliberately: an operator has no API to disable a
single tenant. The incident paths for one misbehaving workspace are: ask the
customer to call chat.tenants.patch(id, {disabled: true}), disable the whole
project with patchProject, or go in with psql (after which the per-instance
tenant cache is stale for up to 15 minutes).
createProject always creates a default tenant in the same transaction — a
project with no tenant could not mint a usable token, since tid is mandatory —
and answers {project, apiSecret, defaultTenant, created}. Replay compares the
project's name/settings against the live row and the default tenant by
normalised slug only: the customer's later renaming or reconfiguring of that
tenant is their business and must not turn the operator's replay into a 409.
A replay after a rotation hands back the current secret, and a replay whose
default tenant the customer has since disabled still succeeds.
rotateSecret, its grace window, and what it takes down
chat.admin.rotateSecret(projectId, {graceHours}) is the one operation the retry
engine refuses outright (retry: 'never'), because a second rotation would
invalidate the secret the first one just handed you. It returns
{project, apiSecret, previousExpiresAt} — no created flag; the previous expiry
is what stands in for one.
The semantics sdk-example-core.int.test.ts proves against a live server:
- The
api_keydoes not change. Rotation is a credential change, not a re-provisioning, sobaseUrlandapiKeyin your config stay put. - The new secret signs tokens the server accepts immediately on the instance
that served the call —
ProjectsServiceinvalidates that instance's project resolver cache, so there is no propagation delay to wait out there. - The old secret keeps working until
previousExpiresAt(token.service.ts).graceHoursis honoured as given (default 24, max 168), which is what lets a fleet mid-deploy run instances on both secrets without an outage. - A secret that was never this project's is rejected, so neither of the two above is an unauthenticated route answering yes to everything.
Roll it out in that order: rotate, deploy the new secret, then let the window expire. Nothing needs to be timed to the second.
Blast radius — read this before rotating. The credential belongs to the
project, so rotation invalidates, after the grace window, the tokens of every
tenant in that project, including tenants the call never named and whose owners
were not told. Guest tokens default to a 7-day TTL, which is longer than the
24-hour default grace, so a rotation cuts off live support-widget sessions across
every tenant mid-conversation. sdk-tenants.int.test.ts pins this by minting for
a second tenant and watching that token die too.
grace_hours: 0 is not instant revocation. The project is cached by api_key
for 15 minutes per instance and the invalidation is local to the instance that
handled the request, so the old secret can still verify on another instance for
up to 15 minutes after a zero-grace rotation. Plan an emergency revocation around
that bound rather than assuming the call is a kill switch.
If a secret leaks, rotating is not sufficient. Since M14 the project's
api_secret also carries control-plane authority over every tenant in the project
(chat.tenants.* needs nothing else). Whoever held the leaked secret could have
disabled all but one tenant, created tenants up to the cap of 100, changed guest
settings, and — the one that keeps working after you rotate — replaced any
tenant's fcm_service_account, pointing that tenant's push notifications at
their own Firebase project. So the runbook is:
rotateSecret(projectId, {graceHours: 0}), and expect up to 15 minutes of residual acceptance on other instances.- Then walk every tenant of the project with
chat.tenants.list()and re-check each one'sdisabledAt, Firebase credential and guest settings against what you expect. Re-set the Firebase credential rather than trusting that it looks unchanged — no route echoes it back, so "looks fine" is not an observation you can make from the API. - Look for tenants you did not create.
Layout
| Path | What it is |
| --- | --- |
| src/generated/ | Generated, committed. Operation table + wire types derived from @chatcore/contracts. Never edit by hand. |
| codegen/ | The generator. emit.ts holds the pure emit functions, generate.ts is the CLI that writes the files. |
| src/index.ts | The closed public surface — an explicit allowlist, locked by test/export-surface.test.ts. |
Rule for the public surface
src/index.ts may not import from src/generated/ or src/internal/, in any
form — export … from, export type … from, or a local import type that is
re-exported afterwards. The guard rejects the module specifier itself, so there
is no wording that gets around it.
That means every publicly exported type lives in a non-internal module (put
them in src/types.ts, or next to the façade class that owns them). A public
type that references an internal type would otherwise force index.ts to reach
into internal/, and the first person to hit that will be tempted to weaken the
guard instead of moving the type. Move the type.
How the secret is protected
api_secret lives only in a closure — never a property, so Reflect.ownKeys
on a ServerTokenCache is empty. test/token.test.ts enforces this statically:
it taints the names that hold secret material, propagates through bindings, and
requires every read to sit in one of six sanctioned positions. Its limits are
listed under documented blind spots in that file; read them before trusting a
green run.
Any new src file that names apiSecret, masterApiKey or api_secret must
be added either to SOURCES (analysed) or to SECRET_BY_CONTRACT (exempt, with
a reason) in that same file. A coverage test fails until it is classified. The
admin surface — facade/namespaces/admin.ts and the admin result types in
facade/resources.ts — is pre-classified as exempt, because it returns a
server-minted api_secret to the caller by design.
facade/namespaces/tenants.ts is deliberately on neither list, and it should
stay that way: a tenant owns no credential (§26.2), so the file names none — not
in code and not in prose. The analyser reads comments too, so an explanatory
sentence about "the project's api_secret" is enough to turn it red. If that
happens, rewrite the sentence; do not add an exemption.
Regenerating src/generated/
pnpm generate reads @chatcore/contracts through its build output, so build
contracts first:
pnpm --filter @chatcore/contracts build
pnpm --filter @chatcore/server-sdk generategenerate runs node codegen/generate.ts directly — Node 24 strips the types,
so there is no build step and no tsx/ts-node dependency. Node detects the
ESM syntax and reparses the file as a module; the resulting
MODULE_TYPELESS_PACKAGE_JSON notice is silenced by a flag in the script,
because the published package is CommonJS and "type": "module" is not an
option. Two consequences show up in the codegen sources: relative imports carry
an explicit .ts extension (legal only under tsconfig.codegen.json, which is
why pnpm typecheck runs two configs), and paths are anchored on
process.argv[1] because __dirname does not exist there.
The generated tables are checked by three different things, because each is blind to what the others catch:
test/generated-freshness.test.tsre-runs the emit functions in memory and compares byte-for-byte with the committed files. Catches a hand edit, or a contracts change nobody regenerated. Blind to a bug in the emitter, which would regenerate a matching-but-wrong artifact.test/generated-projection.test.tscompares the committed tables againstOPERATION_REGISTRYitself — every operation, in order, with every projected field, and fortypes.tsthe full field list and type expression of every entry. This is what catches an emitter bug.pnpm typecheck, which is where the type-level assertions actually run: vitest transpiles without checking types, so a wire type that collapsed toneverwould passpnpm test.
The tests resolve @chatcore/contracts to its source, not dist (see the alias
in vitest.config.ts). pnpm test does not build, so without that a contracts
edit with no rebuild would leave every guard agreeing with a stale registry.
Wire types
src/generated/types.ts maps each operation to its wire shapes. Direction is
not cosmetic:
params,queryandbodyusez.input— what the caller passes. A.default()/.prefault()field is optional there, so callers never have to spell out server-side defaults.responseusesz.output— what comes back, defaults filled in.
Both go through typeof <schema> rather than the hand-written aliases in
contracts, because contracts exports no z.input aliases at all and a few
schemas export their type under a different name than the schema itself.
