@kontourai/relay
v0.7.1
Published
Portable, provider-neutral model invocation contracts, adapters, and conformance tools
Readme
Relay
Relay provides one provider-neutral way to invoke models across direct SDKs, local engines, hosted services, and agent frameworks.
Relay carries an invocation; it does not decide which model to use or what a response means. Bearing owns capability evidence, Datum owns configuration resolution, and domain products retain their prompts and interpretation.
Bearing -> Datum -> Relay -> Traverse / other domain consumersCore contract
import { FakeModelRuntime, checkRuntimeConformance } from "@kontourai/relay";
const runtime = new FakeModelRuntime([{
provider: "fixture",
model: "fixture-1",
modelSource: "configured",
outputText: "",
toolCalls: [{ id: "1", name: "submit", input: { value: 42 } }],
usage: { totalTokens: 10 },
latencyMs: 0,
}]);
const result = await runtime.invoke({
messages: [{ role: "user", content: "Return a structured value." }],
tools: [{ name: "submit", inputSchema: { type: "object" } }],
toolChoice: { type: "tool", name: "submit" },
});
await checkRuntimeConformance(runtime);result.model is the served model only when result.modelSource is
"provider-reported": the provider returned that identity for this invocation.
"configured" means the runtime echoed the model it was configured with, which
may be an alias the provider resolved to something else. The Anthropic-compatible
runtime reports the response's model, the AI SDK bridge reports
response.modelId when the provider model supplies one, the Claude Code harness
reports the served model when its modelUsage names exactly one model, and the
Codex and OpenCode harnesses report configured. Every bundled runtime sets
modelSource and checkRuntimeConformance fails a runtime that does not; the
field is optional in the type for one release so third-party runtimes keep
compiling.
Process-backed harnesses
@kontourai/relay/process provides an abortable, output-bounded transport for
non-interactive model harnesses. A thin harness profile owns request encoding,
response parsing, typed failure classification, and an honest capability
declaration. The transport never interprets prompts, chooses a model, discovers
credentials, or copies constructor environment into results and errors.
import { createProcessRuntime } from "@kontourai/relay/process";
const runtime = createProcessRuntime({
id: "my-harness",
executable: "my-harness",
capabilities: {
structuredTools: false,
streaming: false,
abort: true,
usage: false,
},
codec: {
prepare: (request) => ({ args: ["run", "--json"], stdin: JSON.stringify(request) }),
parse: (output) => normalizeHarnessOutput(output.stdout),
},
});Products should depend on ModelRuntime, not the process profile. This keeps a
single workload definition portable between a locally authenticated harness
and an SDK/API runtime selected later by the host or Dispatch. A profile must
report unsupported capabilities rather than silently approximating them.
Structured-output fidelity is part of that capability declaration. Relay's
built-in SDK and schema-enforced harness adapters report "native"; OpenCode's
explicit prompt-enforced mode reports "prompted"; an adapter without a
structured-output path reports "unavailable". Hosts can therefore apply one
routing policy without branching the workload definition by runtime.
Output-token-limit fidelity is declared separately. "native" means the
runtime receives and enforces maxOutputTokens, "approximated" means the
adapter can only make a best effort, and "unavailable" means the runtime has
no per-invocation hard-limit control. The field is optional so existing custom
runtimes remain source-compatible, while all built-in profiles declare it.
Usage receipts may also include provider-reported cache read/write tokens and
costUsd; Relay preserves those values but never estimates missing cost.
Physical batching is a separately declared capability. A runtime may expose
invokeBatch() only when one call maps to one provider- or runtime-native
physical operation; concurrent invoke() calls do not qualify. Batch outcomes
retain request order and carry typed per-item failures so one rejected item
does not erase successful siblings. The built-in hosted SDK and local harness
profiles currently report physical batching as unavailable. FakeModelRuntime
implements the contract for deterministic consumer and conformance tests.
checkPhysicalBatchConformance() issues one caller-supplied probe batch and
reports only counts and contract checks; it never copies request, response, or
provider-diagnostic content into the report.
Claude Code profile
@kontourai/relay/claude-code uses Claude Code's non-interactive JSON output
and native JSON Schema validation. It supports text-only requests and exactly
one explicitly selected structured tool. Automatic tool choice and required
choice among multiple tools are rejected because the CLI cannot guarantee
those Relay semantics through its structured-output surface.
import { createClaudeCodeRuntime } from "@kontourai/relay/claude-code";
const localRuntime = createClaudeCodeRuntime({ model: "sonnet" });Authentication remains owned by the installed harness. Relay does not inspect, copy, or persist its credential configuration.
Claude Code does not currently expose a per-invocation output-token ceiling
through this profile, so it reports outputTokenLimitFidelity: "unavailable".
When the CLI reports more output tokens than requested, Relay returns
OUTPUT_TOKEN_LIMIT_NOT_ENFORCED in the result warnings. Its JSON receipt's
cache-token and total-cost fields are preserved when present.
Codex profile
@kontourai/relay/codex uses codex exec JSONL events and its native output
schema file. Relay creates the schema and a neutral default working directory
for one invocation, then removes both. A host may provide an explicit working
directory when repository context is intentionally part of the runtime target.
import { createCodexRuntime } from "@kontourai/relay/codex";
const localRuntime = createCodexRuntime({ model: "gpt-5" });Like the Claude Code profile, Codex currently guarantees text or one explicitly selected structured tool. Unsupported tool-selection semantics fail explicitly.
OpenCode profile
@kontourai/relay/opencode consumes opencode run --format json events and
accepts any configured provider/model identifier, including a GLM model.
OpenCode does not currently expose a run-level output-schema flag, so structured
tools are rejected by default.
import { createOpenCodeRuntime } from "@kontourai/relay/opencode";
const localRuntime = createOpenCodeRuntime({ model: "zai/glm-5" });A host may explicitly select structuredOutput: "prompted". That mode reports
structuredToolsFidelity: "prompted", injects the selected schema into the
request, and attaches a warning to successful results. It is intentionally not
presented as equivalent to the native schema enforcement in Claude Code or
Codex, and malformed JSON remains a retryable typed failure.
Descriptions and usage limits in harness profiles
A harness CLI takes the selected tool's schema as an output constraint, not as
a tool definition. All three profiles therefore put the tool description and
the schema's field descriptions into the prompt, so instructions written there
reach the model. Field descriptions are collected from the schema root,
properties, a single-schema items, anyOf/oneOf/allOf branches, and
$defs/definitions. A description under another keyword (prefixItems,
tuple-form items, additionalProperties, patternProperties,
if/then/else, not) is not added to the prompt; the CLI still receives
it inside the schema.
When a CLI reports that its own usage limit or rate limit was hit, the profile
fails with RATE_LIMITED and retryable: true, the same code and flag the API
adapters return for an HTTP 429. The message carries a short reason such as
Claude Code rate limited: weekly limit reached; resets Oct 4. The reason is
assembled from fixed phrases and a date, time, or duration matched by a strict
pattern, read from the line that reported the limit; the CLI's output is never
copied into it, and the reset time is available only as text in the message.
Evidence is weighed in this order, on a nonzero exit and on a failed run that exits zero alike:
- the CLI's structured error report with an authentication status (401 or
403, where the CLI reports one):
AUTHENTICATION_FAILED; - the CLI's structured error report naming a limit or carrying status 429:
RATE_LIMITED, even if stderr holds an unrelated line that looks like an authentication problem; - stderr text that looks like an authentication problem:
AUTHENTICATION_FAILED, even if it also mentions a limit; - stderr text that mentions a limit:
RATE_LIMITED.
Any mention of a rate limit (rate limit, RateLimitError, too many
requests, …) counts. Usage, session, credit, and quota wording counts only in
the forms the CLIs and providers use to say the limit was hit, so a
context-window or turn limit is not reported as a rate limit.
Declarative runtime profiles
Applications can share one PROFILE:MODEL definition without copying adapter
switch statements:
import { createModelRuntimeProfile, parseModelRuntimeProfile } from "@kontourai/relay/runtime-profile";
const runtime = createModelRuntimeProfile({
...parseModelRuntimeProfile("codex:gpt-5"),
cwd: process.cwd(),
});This entrypoint constructs exactly one runtime. It does not select providers,
define fallback, or own budgets; compose multiple runtimes through Dispatch
when the application needs those policies. OpenCode structured tools require
the explicit allowPromptedStructuredOutput option. Hosted credentials remain
explicit constructor inputs and are never read by the profile resolver.
Anthropic-compatible runtime
The optional /anthropic entrypoint loads @anthropic-ai/sdk only when a
client is not injected:
import { createAnthropicRuntime } from "@kontourai/relay/anthropic";
const runtime = createAnthropicRuntime({
model: "claude-sonnet-4-6",
apiKey: process.env.ANTHROPIC_API_KEY,
});One invoke() sends exactly one provider request: the SDK's own retries are off
by default (maxRetries: 0), so the caller or router that records attempts is
the only component retrying. Pass maxRetries to opt back in and timeoutMs to
bound each request (unset keeps the SDK's default). A timeout surfaces as a
retryable PROVIDER_UNAVAILABLE failure. Both options are ignored when a
client is injected, because the caller owns that client's configuration.
createModelRuntimeProfile forwards both options to the anthropic profile.
AI SDK v3 framework adapter
The optional /ai-sdk entrypoint bridges Relay with frameworks that consume
the Vercel AI SDK v3 model contract. Both directions are supported:
import { createAiSdkModel, createAiSdkRuntime } from "@kontourai/relay/ai-sdk";
// Use an AI SDK provider model as a Relay candidate.
const candidateRuntime = createAiSdkRuntime({ model: providerModel });
// Give a Relay runtime— including a policy-backed runtime — to an AI SDK host.
const frameworkModel = createAiSdkModel({ runtime });Text, function tools, tool selection, assistant tool calls, tool results, usage, finish reasons, and abort signals cross the adapter. Provider-defined tools, files, reasoning parts, approval parts, sampling controls, response formats, seeds, and provider-specific options are not portable in the current Relay contract and produce an explicit warning or typed invalid-request error.
Relay v0.2 has no native streaming contract. createAiSdkModel() therefore
buffers one Relay invocation and emits a valid compatibility stream afterward;
the stream includes a compatibility warning and must not be described as live
token streaming.
Credentials are adapter-construction data and must never be placed in Relay requests, results, replay records, or conformance reports. Recording captures request content by design, so hosts must apply their own content-handling policy.
Portable-definition parity
Relay's conformance suite sends the same messages, tool schema, and selected tool through the Claude Code, Codex, OpenCode, and hosted SDK profiles. It compares the normalized tool name and input while allowing transport identity, usage, latency, and warnings to differ. This is the portability guarantee: domain meaning is stable, while the application chooses the runtime binding.
Boundaries
Relay does not own model selection, credential storage, extraction, review, workflow, trust, agent personas, or application tools. See CONTEXT.md.
