@volter/twin-togetherai
v0.1.37
Published
Local Together AI twin — a faithful, stateful local Together API your real `together-ai` SDK talks to unmodified. The model is stubbed (deterministic) and embeddings/rerank/images/transcripts are deterministic stubs, but the protocol envelope (chat comple
Readme
@volter/twin-togetherai
A local, faithful Together AI API twin your real together-ai SDK talks to unmodified — point
the client at the twin's baseURL and chat.completions.create, streaming, models, files,
batches, fine-tunes, embeddings, rerank, the image endpoints and the audio endpoints all
work. Built on the shared @volter/world-core kernel; see
the model for storage and branching; no parallel side store.
import Together from 'together-ai';
import { createTogetheraiTwinServer } from '@volter/twin-togetherai';
const { port } = createTogetheraiTwinServer({});
// together-ai's baseURL is the API ROOT; it appends `/v1/…` (or `/v2/…` for videos) itself.
const client = new Together({ apiKey: 'twin-key', baseURL: `http://127.0.0.1:${port}/v1` });
const res = await client.chat.completions.create({
model: 'meta-llama/Llama-3.3-70B-Instruct-Turbo',
messages: [{ role: 'user', content: 'hello' }],
});
// res.choices[0].message.content starts with "[twin-stub:…]" — a deterministic stub, clearly labeled.
// res.usage carries the three token counts; res.prompt is Together's REQUIRED prompt array.files.upload() needs TOGETHER_API_BASE_URL — the one documented exception
The SDK's custom upload (client.files.upload(path, purpose)) is NOT an ordinary SDK call:
together-ai reads TOGETHER_API_BASE_URL at module load and ignores the client's
baseURL (lib/upload.js:12, [email protected] — it then POSTs
${baseURL}/files?… itself to drive its 302-redirect flow). Unset, that var defaults to
https://api.together.xyz/v1, so files.upload() egresses to the real Together API with the
caller's key even when the client points at the twin. Before the first import 'together-ai'
(the read happens at module load):
TOGETHER_API_BASE_URL=http://127.0.0.1:<port>/v1Every other SDK call rides baseURL (or the injector's api.together.ai interception — a world
wired through TOGETHER_TWIN_URL gets TOGETHER_API_BASE_URL=<twin>/v1 injected from this pack's
descriptor automatically) and needs nothing else. The integration suite pins the same var for the
same reason (togetherai-sdk.integration.test.ts).
Together is not OpenAI — the differences this twin models
Together exposes an OpenAI-compatible inference surface, and treating "compatible" as "identical"
is how a Together twin gets built wrong. The load-bearing differences, each grounded in
[email protected] (Stainless-generated from
Together's own OpenAPI), Together's published spec (141 paths / 156 operations) and Together's
docs, read 2026-09-16:
| | Together | OpenAI |
|---|---|---|
| base URL | https://api.together.ai/v1/… (TOGETHER_BASE_URL) | https://api.openai.com/v1/… |
| error envelope | { error: { message, type, param, code } } — message+type required, param/code nullable | same shape |
| n | 1..128 — an OpenAI-compatible freedom | no 128 cap |
| logprobs | an integer 0..20 (not a boolean) | boolean |
| context overflow | 403 ("context length exceeded"), with context_length_exceeded_behavior: truncate\|error | 400 |
| spending limit | 402 — monthly spending limit reached | — |
| overload | 503 engine_overloaded | — |
| rate limit | 429 with types dynamic_request_limited/dynamic_token_limited + x-ratelimit-reset; success responses carry NO rate-limit headers | x-ratelimit-* on every response |
| finish reasons | adds eos | stop/length/tool_calls/function_call |
| chat response | prompt: [] REQUIRED on the schema; usage is nullable | no prompt array |
| streaming | every chunk carries usage + warnings | usage only in an opt-in tail chunk |
| assistant message | has reasoning / reasoning_content; has no refusal | has refusal |
| files | purposes fine-tune/eval/batch-api, capital-F Processed/FileType, processing_status pipeline; upload is a 302-redirect flow in the SDK or multipart /files/upload | batch/fine-tune/assistants purposes |
| batches | 201 BatchJobWithWarning {job}, list answers a bare array, errors are a bare string, UPPER-CASE statuses, endpoints chat/completions + audio | 200 batch object, {object:'list',data}, cancelling statuses |
| fine-tunes | /v1/fine-tunes (NOT OpenAI's /v1/fine_tuning/jobs), 9-state lowercase enum, delete answers {message} (not {id, deleted}) | different path + shape |
| model ids | slash-namespaced (meta-llama/Llama-3.3-70B-Instruct-Turbo) | flat |
The twin serves only /v1/… (the inference half). The v2 management half (endpoints,
deployments, compute, rl, videos, evaluation, queue, tci) is real Together surface this pack does
not serve — it 404s like any other unmodeled path, never a fake success. Assistants/Threads,
moderations, OpenAI's responses API and OpenAI's batch/file/fine-tuning shapes do not exist at
Together, so the twin refuses them too.
The generative-stub design (be honest)
The twin cannot run the model — there are no weights here. So POST /v1/chat/completions
returns a deterministic STUB completion that is unmistakably a twin stub (it carries a
[twin-stub:<model>] marker and echoes your prompt), POST /v1/embeddings returns deterministic
pseudo-vectors (L2-normalized, seeded from the input hash, at each model's documented
dimensionality), POST /v1/rerank returns deterministic pseudo-scores, and the audio/image
endpoints return deterministic labeled stubs. They never pretend to be real model output.
What is faithful is the entire protocol envelope:
- chat response shape —
{ id, object:'chat.completion', created, model, choices:[{index, message, finish_reason}], prompt:[], usage } - streaming via SSE — real
data: {…chat.completion.chunk with choices[].delta…}chunks ending withdata: [DONE](role chunk → reasoning/content/tool_call deltas → finish_reason chunk → the empty-choices usage tail; every chunk carries Together'susage+warnings; driven through an injected sink — no real sockets in the handler) - tool / function calling —
tool_calls+finish_reason:'tool_calls',tool_choice(none/auto/required/named),parallel_tool_calls, and the deprecatedfunctionsarray answering the deprecatedfunction_callshape - reasoning —
reasoning_effortlow/medium/highsurfaces thereasoningfield on the assistant message - structured outputs —
response_formatjson_object/json_schema(Together REQUIRESjson_schema.name, ≤64 chars[a-zA-Z0-9_-]— enforced) n(1..128),max_tokens,stop,seed,echo(the REQUIREDpromptarray), integerlogprobs(0..20),compliance:'hipaa'; the OpenAI fields Together accepts-but-ignores (service_tier,store,metadata,prediction) change nothing; deterministic usage- Stub honesty: the deterministic completion is ONE continuation regardless of
n(every choice repeats it);stopsequences are honored deterministically — the text is cut at the earliest occurrence of any stop string (the envelope validates and echoes them).
The genuinely stateful + static surface is real, not a stub:
GET /v1/models(+/models/:id) — static catalog answering a bare array ofModelInfowith Together'stype/display_name/context_lengthGET /v1/whoami— static identity for the presented key- Files —
POST /v1/files/upload(the spec's create) +GET/DELETE /v1/files{,/:id,/content}(+ the SDK's 302-redirect flow withx-together-file-id), kernel-backed, Together's closed purpose/type sets,processing_statuson fine-tune files - Batches —
POST/GET /v1/batches(+ cancel →CANCELLED), 201BatchJobWithWarning, bare-array list, closed 3-endpoint set. The asynchronousVALIDATING→IN_PROGRESS→COMPLETEDprogression is a filedtodo, not modeled: doing it on a read made a GET write to the log (breaking the read-only contract) and minted an action with no vendor endpoint to push it to. - Fine-tunes —
POST/GET/DELETE /v1/fine-tunes(+ cancel, events, checkpoints, estimate-price, preview, models/limits), purpose-checked training files; delete answers{message}and estimate-price the SDK'sestimation_availableunion
Vendor-shaped errors throughout: 400 / 401 / 402 / 403 / 404 / 405 (read-only guard) /
429 / 503. Every envelope carries the required message + type and only keys Together's
ErrorData declares.
Rate budget (live calls)
liveTogetheraiExecute — the ONE place this pack issues a real api.together.xyz request — routes
every call through the kernel's fail-closed RateBudget. Together publishes NO fixed request limit
(dynamic, per organization per model), so the declaration pins the kernel's fallback ceiling
(60 units / 60s, defaultWeight 2 = 30 calls/min) and buys resolution DOWNWARD: inference endpoints
cost 6, so at most 10 land in a window. A 429's x-ratelimit-reset becomes a persisted cooldown.
There is no option to disable the guard.
Scenario scripting (twin-only)
Chat completions can be scripted by MSW-shaped handlers in a world dir's handlers/togetherai.json,
run by the ONE kernel scenario engine with this pack's adapter:
{ "handlers": [
{ "on": { "userTextIncludes": "refund", "hasTool": "create_job" },
"respond": { "toolCalls": { "name": "create_job", "arguments": { "title": "Refund order" } } }, "once": true },
{ "on": { "reasoningEffortEquals": "high" }, "respond": { "error": { "type": "engine_overloaded" } } }
] }Load it with --scenario FILE / createTogetheraiTwinServer({ scenarioPath }) /
TWIN_TOGETHERAI_SCENARIO. The read doors GET /twin and GET /twin/scenario explain and inspect
it. A malformed file throws at load, never a silent fall-back. Scenario support is twin-only
scaffolding for eval worlds, not vendor surface, so it is deliberately not in the capability
manifest.
CLI
world-togetherai serve [--port N] [--root DIR] [--read-only] [--scenario FILE]
world-togetherai conformance [--root DIR]Coverage
The capability manifest (togetherai-capabilities.ts) is authored against Together's published
OpenAPI spec — 141 paths / 156 operations — with every entry done (with a runnable verify) or
todo. The v2 management half is filed honestly as todo at niche tier. Run
world-togetherai conformance for the offline protocol checks.
