ns-kiro-core
v0.4.0
Published
Shared Kiro (AWS CodeWhisperer/Q) protocol core — credentials, model catalog, streaming; a dependency, not for direct install
Readme
ns-kiro-core
The Kiro (AWS CodeWhisperer/Q) protocol, with no host types in it.
This package is the Kiro vendor core of ns-bridge: everything a Kiro client needs that is not specific to one coding agent.
Published on npm as ns-kiro-core. Pulled in as a dependency by the Kiro host
adapters — @ngosangns/ns-pi-provider, ns-omp-provider-kiro,
ns-dsh-llm-kiro — not meant to be installed on its own.
- Endpoints — SSO region to management/runtime host resolution.
- Model catalog — the bootstrap list, the authenticated regional catalog,
and a validated on-disk cache at
~/.ns-kiro-provider-models-cache.json. Per-model billing weights (rateMultiplier) come from kiro-cli, the only source that publishes them. - Usage — Kiro reports no token counts, only a billed amount, surfaced as
usage.credits.usage.inputis derived from the context-usage frame andusage.outputfrom a tiktoken estimate. - Credentials — reads the kiro-cli SQLite store and the Kiro IDE token file,
refreshes IDC / desktop / external-IdP / API-key sessions, and writes
refreshes back so kiro-cli stays on the same token. Both stores can be signed
into different IdC users of the same Kiro profile, so the kiro-cli session is
used and refreshed first, and the IDE's only when kiro-cli holds nothing;
KIRO_AUTH_SOURCE=idereverses that for a machine whose IDE login is the one to use. - Streaming — AWS event-stream framing, thinking-tag parsing, native and text-dialect tool-call recovery, history validation and repair, and the whole retry ladder (transport timeouts, capacity pressure, request-rate windows, 403 credential rotation, degenerate 200s).
Engine
Since 0.4.0 the public entry points (streamKiro, updateKiroModelsCache, fetchKiroUsage, refreshKiroToken) run in the
ns-bridge Go sidecar
when a binary is installed (ns-bridge-bin,
NS_BRIDGE_BIN, or ns-bridge on PATH), and in this package's TypeScript
otherwise. NS_BRIDGE_ENGINE=ts (or NS_BRIDGE_ENGINE_KIRO=ts) keeps
everything in-process; …InProcess exports always do. Per-process state
(profile-ARN and region caches, the cache-read estimate, catalog refresh scheduling) stays in TypeScript and is carried across each call.
The TypeScript implementation is now the fallback. A later major release is expected to slim this package to the facade, types and credential stores, with the protocol living in the binary only; nothing changes for callers of the exports above.
The neutral seam
streamKiro takes a request built from this package's own vocabulary and yields
its own events:
import { streamKiro, resolveKiroCredentials, getCachedModels } from "ns-kiro-core";
const credentials = await resolveKiroCredentials();
const model = getCachedModels("us-east-1").find((m) => m.id === "claude-sonnet-4-6");
for await (const event of streamKiro({
model: { ...model, region: "us-east-1" },
messages: [{ role: "user", content: [{ type: "text", text: "hello" }] }],
accessToken: credentials.access,
effort: "medium",
})) {
if (event.type === "text_delta") process.stdout.write(event.delta);
}An adapter owes two translations — host messages in, host stream events out —
and nothing else; for Pi-family hosts and the Harness, ns-bridge-core already
does both. loginKiroFromSession / resolveKiroRequestCredentials (session
login and per-request credentials) and toKiroModelForHost (catalog entry to a
host model) are the other pieces every host needs, so they live here too. Block indexes are monotonic across the whole response,
including across an internal retry, so a host that cannot un-deliver a block
still receives a coherent sequence; canDiscardEmittedBlocks tells the core
which kind of host it is talking to.
The stages underneath
streamKiro is orchestration over four pieces, each exported for use on its
own — building a request without sending it, or assembling blocks from events
sourced elsewhere:
| Export | Does |
| --- | --- |
| buildKiroRequest | Neutral messages to a wire request: history shaping, tool specs, pre-send repair. Pure, no I/O |
| readKiroEventStream | AWS event-stream framing and the stall timeouts; yields parsed wire events |
| KiroResponseAssembler | Content blocks, thinking, tool calls, text-dialect recovery, usage, stop reason |
| abortableDelay, createResponseHeaderDeadline, logCapacityEvent | Transport timing |
What Kiro reports, and what it does not
Measured 2026-09-06 against claude-sonnet-5 in us-east-1. Recorded here so
the questions are not re-opened from first principles.
Token counts: none. Kiro's usage frame is a billing record —
{unit: "credit", usage: 0.0659} — not token counts. usage.input is therefore
derived from the contextUsagePercentage frame, and usage.output from a
tiktoken estimate over what the model emitted. usage.credits carries the
figure Kiro actually bills.
Per-token prices: none. Kiro bills in credits and publishes no per-token
rates, so every model's cost stays zero. kiro-cli chat --list-models does
publish a relative billing weight, which the catalog picks up as
rateMultiplier — 2.2 for claude-opus-5 against 0.05 for qwen3-coder-next.
It is absent when kiro-cli is not installed.
Prompt caching: real, but not controllable. Kiro caches prompts server-side
on its own: a repeated prefix billed ~0.035 credits against ~0.066 for a fresh
one, and a changed prefix went straight back to the full price. There is no way
to ask for it — every model's additionalModelRequestFieldsSchema sets
"additionalProperties": false and allows only thinking/output_config/
max_tokens (Claude) or reasoning (GPT), so a cachePoint or cache_control
field is rejected rather than honoured. Kiro also reports no cache token counts.
Reasoning effort is part of the cache key: changing it misses even when the prompt is byte-identical, and each effort level then warms its own entry.
Stop reasons: reported when Kiro sends one, inferred otherwise. Kiro now
closes a turn with a metadataEvent stopReason (END_TURN on every ordinary
turn checked, 2026-10-08). MAX_TOKENS is reported as length;
CONTENT_FILTERED, MODEL_CONTEXT_WINDOW_EXCEEDED (phrased
context_length_exceeded) and PAUSE_TURN end the call with an error and no
retry. Without a stop reason, a turn with no tool call that never carried a
contextUsagePercentage frame is reported as length. That frame arrived in
every case checked — a short reply, a ~5000-character one, a tool-call turn, a
model with no effort schema, and a non-Claude model — so its absence does mark
an abnormal turn rather than a normal short answer.
Set KIRO_DEBUG=1 to log the frames verbatim (~/.ns-kiro-provider/logs/) if
any of this needs re-checking against a newer Kiro.
Credit
Ported from pi-provider-kiro by Mike O'Brien (MIT). The port replaces pi-ai's message and event types with the neutral vocabulary above; the protocol behaviour is upstream's.
