kiro-mem
v3.0.0
Published
Persistent cross-session memory system for Kiro CLI. Captures every turn, compresses via Kiro CLI ACP, and injects a compact memory index into later sessions.
Maintainers
Readme
Persistent memory system for Kiro CLI.
Kiro CLI only. Not compatible with Kiro IDE.
Quick Start • How It Works • MCP Tools • Web Viewer • Configuration • CLI • Limitations • License
kiro-mem automatically captures each turn (prompt → tool calls → stop) during Kiro sessions and compresses it into a single immutable Observation — one per closed turn. It never rewrites history or guesses topics up front. At session start it injects a compact index of prior work in the current workspace; the agent scans that menu, then searches and pulls full details on demand. Organizing relevance happens at read time, driven by the agent, not by fragile write-time clustering.
Key Features
- 🧠 Persistent Memory — Keep project context across sessions
- 🧩 Atomic & Immutable — One closed turn → one Observation; text is never rewritten or merged
- 🤖 ACP-native — Compression runs through Kiro CLI ACP — no LLM API key required
- 🔍 Hybrid Search — FTS5 full-text search + local semantic recall, fused with RRF and hard-scoped per workspace; a query with no shared wording can still find the record
- 📊 Read-time Organization — Inject a compact index; the agent pulls relevance on demand (
search→timeline→get_observations) - 🔧 MCP Tools —
search,timeline,get_observations,pin - 🖥 Web Viewer — Local browser UI: browse each Observation next to the turn it came from, preview the exact context the next session will receive, and permanently delete a memory together with its source turn
- 🔒 Privacy Control — Use
<private>tags to redact sensitive content before storage - 🚀 Async Processing — Persistent job queue, no tool-call blocking
- 🔄 Process Keepalive — Worker managed by
launchdorsystemd - 🌐 i18n —
zhandenfor the CLI, the runtime compressor prompt and the Web Viewer UI
Quick Start
Requires Bun and a Kiro CLI build that supports the acp subcommand.
V3 is a clean break from V2, and the only supported upgrade path is a full wipe and reinstall. There is no database migration and no compatibility mode; running a V3 install over V2 data is not supported and is not tested.
kiro-mem stop kiro-mem uninstall --purge # ⚠️ see the warning below npm i -g kiro-mem@3 kiro-mem install⚠️
--purgepermanently deletes all memory and configurationIt removes
~/.kiro-memin full: every captured turn, every Observation, every vector, the job queue, yourconfig.json, and the local Worker token. This cannot be undone and there is no export. If any of that history matters to you, copy~/.kiro-memsomewhere else before running it — note that a V2 copy cannot be read back by V3.
kiro-mem stopfirst is not optional: it prevents the old Worker from writing into a directory that is being removed underneath it.
For a first-time installation:
npm i -g kiro-mem@3
kiro-mem installThe installer checks Kiro CLI ACP availability, creates a high-entropy local Worker token with owner-only permissions, copies the bundled embedding model (~23 MB) into ~/.kiro-mem/models/, lays out the isolated kiro-runtime (compressor sub-agent + prompt), then registers and starts the Worker under launchd/systemd. No API key required — memory compression runs through kiro-cli acp against your existing Kiro session.
The Worker fails fast on startup if kiro-runtime is incomplete (missing agent file, missing prompt, or tools accidentally non-empty), so a broken layout never silently degrades compression purity. Re-run kiro-mem install to repair.
Set As Default Agent
kiro-cli settings chat.defaultAgent kiro-memOr switch inside a chat session:
/agent kiro-memA mid-session switch does not inject the memory index.
agentSpawnis a session-start trigger — its array-format alias is literallySessionStart— and Kiro CLI does not replay it when you switch agents. Measured on kiro-cli 2.19.1 with a probe agent whoseagentSpawnhook appends to a file: starting with--agent probewrites one line, while starting on another agent and then running/agent probewrites none, even though the TUI confirmsAgent changed to probe.Everything else about the switched-in agent is live:
/mcpshowskiro-mem ● running 4 toolsand theuserPromptSubmit/postToolUse/stophooks do fire, so the turn is still captured and compressed — you only lose the injected index for that session. To get it anyway, either start the session on kiro-mem, ask the agent to@kiro-mem/searchexplicitly, or paste the exact string the hook would have injected:curl -s -H "Authorization: Bearer $(cat ~/.kiro-mem/.token)" \ "http://127.0.0.1:37778/context/bootstrap?cwd=$PWD"
Verify Installation
kiro-mem diagnose
kiro-mem status
curl http://127.0.0.1:37778/healthHow It Works
Architecture
- Core Truth Layer —
session_id→turns→turn_events(append-only raw hook payloads) +turn_artifacts(deterministic extraction: tools, files, commands, test/build/lint signals, errors). This layer is never rewritten and can rebuild any projection. - Synthesis Layer — Persistent jobs (
summarize_turn→embed_observation) drive an ACP runtime pool: each prompt goes to akiro-cli acpsub-process running an isolatedkiro-mem-compressorsub-agent declared withtools: []. Any tool-call notification on that session is treated as contamination and the runtime slot is recycled.summarize_turnconsumes the user prompt, the assistant's final response (assistant_response), and the deterministic artifacts, then writes exactly one immutable Observation per closed turn (idempotent onturn_id). It never clusters, merges, or supersedes. On compression failure it degrades to aquality=fallbackObservation carrying only deterministic evidence — never a fabricated result. - Retrieval Layer —
observations_fts(FTS5) and local semantic recall run as two independent legs, fused with RRF and hard-scoped by workspace → MCP tools → context injection. The semantic leg does not need a keyword hit to run: it can recall a record nothing in the query literally matches, bounded by a cosine floor and a semantic-only cap.
At session start, kiro-mem injects a compact Observation index for the current workspace: pinned observations, a few recent entries with a short outcome snippet, and a longer recent index where each line carries an approximate read cost. It is a menu, not the content — no LLM synthesis, no topics. The agent then pulls relevance on demand via search → timeline → get_observations.
Data Model
session_refs— Session isolation metadataturns+turn_events— Per-turn lifecycle row + append-only raw hook payloads (truth layer)turn_artifacts— Deterministic extraction (tools, files, commands, test/build/lint signals, errors)observations— Immutable per-turn memory unit: one closed turn → one Observation, text never rewrittenobservations_fts(FTS5 trigram) +observation_embeddings— Hybrid search backing tablesjobs— Persistent async task queue
MCP Tools
| Tool | Purpose |
| ------------------ | ------------------------------------------------------------------------------ |
| search | Hybrid search observations with type, days, repo/cwd filters |
| timeline | Show temporally adjacent observations around one, in real source-turn order |
| get_observations | Fetch full observation details by ID, with source-turn metadata |
| pin | Mark or unmark an observation for later context |
@kiro-mem/search query="auth module bug" type="bugfix" limit=10
@kiro-mem/timeline observation_id=42 before=3 after=3 mode="scope"
@kiro-mem/get_observations ids=[42,56]search is hard-scoped to the current workspace via computeScopeKey(repo, cwd) — pass repo/cwd to target a specific scope, or all_scopes: true to browse everything. timeline anchors on an observation's source turn (not its ID), so neighbors reflect the real work timeline; use mode="session" to stay within one conversation instead of the whole workspace.
Observation Types (memory_type): decision | bugfix | feature | refactor | discovery | change
Privacy
Use <private> tags to redact sensitive content before storage:
<private>database password is xxx</private>
Help me configure the connectionContent inside <private> tags is replaced with [REDACTED] before it is written to storage — this covers the user prompt, tool payloads, and the assistant's final response, so private text never reaches an Observation or the injected index.
Web Viewer
kiro-mem viewerOpens a local browser UI served by the Worker itself — no CDN, no dev server, no network beyond loopback. It starts on All workspaces; the workspace picker narrows the feed to one project and lists each workspace with its full path, Observation count and last activity. A standing notice remains visible while the global view is selected.
The whole UI renders in one language, chosen by config.language and delivered in the bootstrap response — no bilingual labels. A kiro-mem config change takes effect on the next page load. If the field is absent (an older Worker), it falls back to English.
What it answers:
- What was remembered — a live Feed of Observations, newest first, with type, quality, pin state and a summary/outcome snippet. New Observations appear as their compression commits.
- Whether the memory is faithful — each card opens a side-by-side detail view: the generated fields (
title,summary,request,outcome,learned,next_steps, evidence, concepts, files) next to the Truth Layer they came from (prompt_text, deterministic artifacts, event counts and original byte sizes). Raw event payloads load on demand and are bounded per event and per response, so a 4MB turn cannot stall the page. - What the next session will receive — the context preview renders the exact string the
agentSpawnhook injects, withusedBytes / effectiveMaxBytesand a per-section byte breakdown (frame, trust boundary, usage, pinned, recent-detail, recent-index). The budget is adjustable for preview and clamped server-side to 9500 bytes. - Why a search matched — keyword search over the current scope, with
match_sourceshown per result. Viewer search is keyword-only: it never fabricates asemantic_query_enon your behalf, so results labelledsemanticcome from the agent-facing path, not from this UI. - Whether retrieval is healthy — a panel projecting the same trailing-24h
search_24hcounters and retrieval profile that/healthreports, plus a worker error log drawer.
The Viewer has exactly two write actions, and their confirmation matches their reversibility: pin toggles optimistically and reports the result in a transient toast (including a rollback notice if the Worker refuses it), while delete always opens a modal stating what will be destroyed.
Permanent Deletion
The trash icon on a card (or in the detail header) opens a confirmation showing the Observation id and title, the workspace, the turn's start/stop time, and how many raw events and original bytes will be destroyed. Confirming deletes, in one SQLite transaction:
the Observation → its vectors and semantic-normalization row → its FTS entry → the source turn's turn_artifacts, turn_events and turns row → every pending or terminal job attached to that turn or Observation.
There is deliberately no "delete the memory but keep the turn" mode: a kept turn is exactly what kiro-mem repair re-queues, so the record would come back. After a successful delete the turn cannot be reached by search, timeline, get_observations, context injection or repair, and every open Viewer removes the card immediately.
Refusals are honest rather than partial:
- a related job in
leasedstate returns 409 and changes nothing — an ACP or CPU task cannot be cancelled mid-flight, so retry after it finishes; - any SQL failure rolls the whole transaction back, leaving row counts unchanged;
pindoes not block deletion; it only affects display and injection priority.
This is product-level permanent deletion — not queryable, not retrievable, not rebuildable. It is not a forensic wipe: ordinary SQLite DELETE leaves bytes in the WAL, the freelist, filesystem snapshots and any backup you took, and kiro-mem deliberately does not run an automatic VACUUM or rewrite the database.
Security Model
- The Viewer runs on the same loopback-only Worker and reuses the existing local Bearer token.
/uiand its two static assets are the public shell; every/api/viewer/*route fails closed without the token. kiro-mem viewerhands the token over in the URL fragment, which browsers never send to the server. The page moves it into that tab'ssessionStorageand rewrites the URL, so the token stays out of history, Referer headers and the Worker's logs. Closing the tab ends the session; a 401 shows "runkiro-mem vieweragain" instead of reconnecting in a loop.- Requests must carry an exact loopback
Hostmatching the Worker's port, which blocks DNS rebinding.DELETEadditionally requires anOriginidentical to the page's own origin. - Responses set
nosniff,DENYframing,no-referrer, same-origin COOP/CORP and a CSP ofdefault-src 'none'with only'self'scripts and styles. - Recorded memory is untrusted input and is rendered as text nodes only. There is no
dangerouslySetInnerHTML, no Markdown rendering and no auto-linking anywhere in the bundle.
If the Viewer bundle is missing (a source checkout that has not run bun run build:ui), the Worker still starts normally and /ui returns a readable build hint.
Configuration
Edit ~/.kiro-mem/config.json, or run kiro-mem config for interactive setup:
{
"language": "zh",
"compression": {
"concurrency": 3,
"minWarmRuntimes": 1,
"idleTtlMs": 600000,
"timeoutMs": 30000,
"maxRetries": 2
},
"context": {
"maxOutputBytes": 8192
},
"filter": {
"skipTools": ["introspect", "todo_list", "@kiro-mem/*"]
},
"retrieval": {
"semanticDiscovery": true
},
"runtime": {
"kiroHome": ""
}
}language:zhoren. Drives the CLI output, the runtime compressor prompt and the Web Viewer UI, which renders in this language only.compression.concurrency: number of parallelkiro-cli acpruntime processes (default3). This is a hard ceiling on processes, not just on pool slots: a runtime being shut down keeps its place in the budget until its process is actually gone, so a replacement is never started alongside it.compression.minWarmRuntimes: how many runtimes idle reclamation may never take away (default1, clamped to[0, concurrency]). It is a floor, not a target — a Worker that has never compressed anything still holds zero.0lets the pool empty out completely between bursts, at the cost of an ACP cold start on the next turn.compression.idleTtlMs: how long a released runtime may sit idle before it is retired (default600000, 10 minutes; clamped to[1000, 86400000]when set viakiro-mem config).0turns idle reclamation off — it never means "kill immediately". Negative,NaNandInfinityfall back to the default.compression.timeoutMs: per-prompt timeout in milliseconds, clamped to[5000, 60000]when set viakiro-mem config(default30000).compression.maxRetries: how many JSON-repair retries to attempt before degrading to aquality=fallbackObservation (default2).runtime.kiroHome: isolatedKIRO_HOMEfor the compressor sub-agent. Empty falls back to<dataDir>/kiro-runtime, which is the layoutkiro-mem installlays down.context.maxOutputBytes: byte budget for the injected Observation index, kept below theagentSpawn10KB limit (default8192).retrieval.semanticDiscovery: whether semantic similarity may surface records that keyword search never matched (defaulttrue). See below.
ACP Process Governance
Compression runs in kiro-cli acp sub-processes, and each one costs real resident memory — measured on a dev machine at ~38MB per runtime (9.8MB in the direct child plus a 28.1MB child of its own). Several Kiro windows do not get a Worker each: they all reach the same ~/.kiro-mem Worker over loopback and share one pool, so the ceiling below is a ceiling for the whole machine, not per window.
The pool starts empty and stays empty until the first compression job. From then on:
- a finished runtime is reused for the next job rather than restarted;
- a runtime idle for longer than
idleTtlMsis retired — its process is killed and not replaced — as long as that leaves at leastminWarmRuntimesbehind; concurrencybounds processes, not bookkeeping: a runtime being retired keeps its slot in the budget until its process has actually exited, so a replacement is never spawned next to one that is still shutting down;- a runtime that hit
maxJobsPerProcessor an ACP error/contamination is restarted in place. The slot count does not change, so the warm floor does not block it — even for the last remaining runtime; - a slot whose process died on its own is dropped unconditionally, ignoring both the TTL and the warm floor. A dead runtime is not warm.
Nothing interrupts work in flight: idle time only starts counting after a job releases its runtime, and a busy runtime is never retired.
Two readings tell you whether this is working:
curl -s http://127.0.0.1:37778/health | jq '.acp' # idleRecycles, config, per-slot idleMs
kiro-mem diagnose # same signals, formatted/health.acp separates the three reasons a runtime was closed — jobLimitRecycles, errorRecycles, idleRecycles — plus deadDrops for processes that vanished on their own. restarts keeps its original meaning of in-place restarts only (jobLimitRecycles + errorRecycles), so idle reclamation never shows up as instability. config reports the values in effect in that Worker process, which is not necessarily what config.json says right now: pool parameters are read at Worker start, so an edit applies after kiro-mem stop && kiro-mem start.
On shutdown the pool is closed before the Worker waits for in-flight jobs to record their final state, so no ACP child outlives its Worker. kiro-mem stop, uninstall and uninstall --purge all go through that path, and each only ever signals the PID recorded in .worker.pid — kiro-mem never scans for or kills processes it did not start.
Semantic Discovery and Rollback
retrieval.semanticDiscovery: true is the current code default: a search that carries a legal semantic_query_en runs the semantic leg even when FTS matched nothing in this workspace, so a question phrased entirely differently from the record can still find it. This profile has not passed the final P6 false-recall gate (C1 measured 1.900 against a ≤ 1 bar). Shipping it on by default is a recorded, informed decision — benchmark/reports/release-decision-2026-08-14.md — not a passed gate: turning it off does not mean "two fewer results", it means a search returns nothing at all when no keyword matched, and the blind audit measured that capability as old-record hit@5 going from 0% to 80.77%. The gate reading itself is unchanged, and so is every cost listed in Limitations. A later round scanned floor 0.197–0.325 × cap 1–2 on an independent calibration set and found no safe working point at all — see "No cosine threshold separates the two sides" in Limitations for the readings and, importantly, for which datasets they came from and why neither can be used to pick a threshold. Treat 0.197 as the calibrated default, not as a knob a higher value would make safer. Its current bounds, selected by an earlier full-pipeline parameter scan, are a cosine floor of 0.197 and at most 2 semantic-only results per search. Results that only the semantic leg found are labelled match_source: "semantic" and should be verified before you act on them.
The cap bounds one source label, not the whole page: fts and hybrid results were never subject to it, and once the default time window became unbounded the keyword leg's reach grew with it. Treat the cap as a quota on unverified pure-semantic leads, not as complete page safety.
Setting it to false is the rollback. For queries that carry a legal semantic_query_en it restores the previous behavior — keyword anchor required, floor 0.2, no semantic-only cap, recency tie-break — and takes effect for Kiro sessions started after the edit.
One thing the rollback deliberately does not restore: a search with a missing or refused semantic_query_en stays keyword-only under either setting. The old build scored the raw-v1 space in that case, and that space has no defensible threshold, so a single keyword hit was enough to put an unrelated record on the page labelled semantic. That is a fixed safety hole, not a rollback gap.
# roll back
kiro-mem config --show # shows the active profile
# edit ~/.kiro-mem/config.json: "retrieval": { "semanticDiscovery": false }The switch changes retrieval only. It does not touch the database schema, the stored vectors or the embedding protocol, so it can be flipped back and forth with no rebuild and no data loss.
Retrieval Metrics
curl http://127.0.0.1:37778/health and kiro-mem diagnose report a trailing 24h window of search counters. They hold counts, enum reasons and latency only — never query text, Observation text or workspace paths.
| Field (search_24h) | Meaning |
| --- | --- |
| requests / latencyMsP50 / latencyMsP95 | Search volume and latency. |
| protocolSemanticEn / semanticEnRate | How many searches actually ran in the English vector space. Independent semantic recall only works there, so a low rate means the feature is shipped but mostly unreachable. |
| semanticQueryIssues | Why the English form was unusable, by reason: missing (the agent never passed semantic_query_en) plus guardrail rejections such as untranslated or placeholder. |
| ftsOnly / degradeRate | Searches that fell back to keyword-only because the query embedding was unavailable (Worker down, timeout). A degraded request is a capability failure, not a policy choice. |
| zeroFts / zeroFtsRecalled | Searches with no keyword match at all, and how many of those still returned something via semantic recall. This is the number the feature exists to move. |
| semanticOnlyTotal / semanticOnlyPerRequest / semanticOnlyMax | Volume of unverified semantic-only leads. semanticOnlyMax must never exceed the cap. |
| comparableVectors (avg) | How many stored vectors the semantic leg scored. The candidate pool is scope-wide, so this tracks how much of the workspace a search actually compared against rather than saturating at a constant. |
| scopeVectors (avg / min / measured) | Vectors in the active space across the WHOLE searched scope — not the candidate pool. This is the number that answers "has this workspace been embedded under this protocol?". null (not counted) when the semantic step never ran, which is deliberately different from 0. |
| emptyScopeRequests | Searches whose scope had no vectors at all. Non-zero means semantic recall was structurally impossible for them — a rebuild or job-backlog problem, not a relevance one. |
/health also reports the active profile under retrieval, read from disk on every call, so "did my rollback take effect?" is answerable without reading code. It reports what the next session will serve — a session already running keeps the profile it started with.
CLI
kiro-mem install
kiro-mem status
kiro-mem start
kiro-mem stop
kiro-mem config
kiro-mem config --show
kiro-mem diagnose
kiro-mem repair
kiro-mem viewer
kiro-mem uninstall
kiro-mem uninstall --purgeSystem Requirements
- Bun: Latest version
- Kiro CLI: >= 2.3.0 — earlier versions lack the
acpsubcommand and theKIRO_HOMEoverride that the isolated compressor sub-agent depends on. Runkiro-cli --versionandkiro-cli acp --helpto verify. - macOS / Linux: Required for Worker keepalive via
launchd/systemd
Limitations
| Limitation | Impact | Mitigation |
| ----------------------------------- | --------------------------------------------------------------------- | --------------------------------------------- |
| Capture is best-effort | Hooks give the Worker ~700ms and never block your turn, so a Worker restart or transient error can drop a raw event. Dropped input cannot be reconstructed later — memory is not guaranteed to be complete | Misses are counted per hook and reason; see capture_misses_24h in /health and the warning line in kiro-mem diagnose |
| Deletion is one record at a time, and manual | Nothing expires on its own: there is no retention policy, no automatic cleanup and no batch prune, because kiro-mem does not decide which of your memories are worthless. The Web Viewer deletes one Observation together with its source turn per confirmation, and there is no export, no bulk delete, no soft delete and no undo. A leased related job makes that one deletion return 409 until it finishes | Delete from the Viewer for anything you want gone; <private> tags keep sensitive content out of storage in the first place; kiro-mem uninstall --purge still wipes everything at once |
| Deletion is not a forensic wipe | Ordinary SQLite DELETE leaves the old bytes reachable in the WAL, the freelist, filesystem snapshots and any backup you already took. kiro-mem deliberately runs no automatic VACUUM and never rewrites the database, so "permanently deleted" means unreachable to every kiro-mem read path — not scrubbed from the disk | Treat the promise as product-level: not queryable, not retrievable, not rebuildable. For media-level guarantees use full-disk encryption and control your own backups |
| Raw event payloads are capped | A tool response over 32KB per string field is truncated, and a single turn stores at most 4MB of raw payload. Truncation is marked inline, but the dropped bytes are not recoverable | payload_size still records the original size; artifacts extraction keeps working on the capped payload |
| Requires Kiro CLI ACP | Compression cannot run without a working kiro-cli acp subcommand | kiro-mem diagnose runs an ACP smoke test |
| One warm ACP runtime stays resident | Between bursts the pool keeps minWarmRuntimes processes alive so the next turn does not pay ACP cold start — ~38MB at the default of 1. Idle reclamation never drops below that floor, and a second Worker on an isolated KIRO_MEMORY_DATA_DIR keeps its own floor rather than sharing one; ordinary multi-window use does not create a second Worker | Set compression.minWarmRuntimes: 0 to let the pool empty out, or shorten compression.idleTtlMs. Both trade memory for cold-start latency on the next compression |
| Pool parameters are read at Worker start | Editing concurrency, minWarmRuntimes or idleTtlMs in config.json does not affect a Worker that is already running, so process behavior and the file can disagree until a restart | /health.acp.config reports the values in effect in that process; kiro-mem stop && kiro-mem start (or kiro-mem config, which restarts for you) applies the edit |
| agentSpawn output limit 10KB | Injected index must stay compact | Budget-controlled context builder |
| Mid-session /agent switch injects nothing | agentSpawn is a session-start trigger, so switching into kiro-mem with /agent leaves that session without the memory index. Capture and the MCP tools do keep working — only the injected menu is missing | Start the session on kiro-mem (--agent kiro-mem or chat.defaultAgent), ask for @kiro-mem/search explicitly, or paste /context/bootstrap output — see Set As Default Agent |
| Search queries shorter than 3 chars | Falls back to LIKE, less precise | Use longer terms when possible |
| Search has no default time limit, and no retention policy | search defaults to the whole history, so the searchable range equals what the injected index can show — a record from two years ago is reachable. The cost is that the semantic leg's working set grows with the corpus and nothing ever expires. Measured ceiling: 50,000 records spread over 3 years, p95 203–260ms and search-loop memory +264…+294MB across three runs. Beyond that size the behavior is unmeasured | Pass days=N to narrow one search. retrieval.semanticDiscovery: false restores the previous profile, which also restores its 200-record pool |
| Semantic-only results are unverified | A query with no keyword overlap can now find records through meaning alone, but those results rest on vector similarity only, and a query about work this project never did can still return up to 2 plausible-looking records. An independent blind audit measured a mean of 1.9 returned records across 30 hard negatives — the cap is reached on nearly every one of them | They are labelled match_source: "semantic" and capped at 2 per search; verify with get_observations or the current code before acting. The cap covers only the semantic label — fts and hybrid results are not bounded by it |
| No cosine threshold separates the two sides | Two measurements, on two datasets that are both already consumed — neither is a new independent validation set. P3 set (40 records / 131 queries; independent of the earlier blind-audit set, but consumed by P3's own floor × cap selection): the best-scoring false lead reached cosine 0.604 while the best-scoring true answer reached 0.602 — the false lead outscored the real one. So no single absolute cosine floor can both zero the 17 zero-FTS foreign-domain negatives and keep every primary_gold semantically reachable; a floor above 0.604 costs all 61 of them that reachability. P4.0 probe (a different, also-consumed set: 26 relevance + 30 hard-negative queries): relative normalization measured worse than the raw score — z-score AUC 0.735 vs raw top-1 AUC 0.841 — because a query about work that never happened has a lower corpus background, so normalization rewards it | Read the boundaries with the numbers. Both sets are consumed: they are descriptive evidence only, and no threshold may be selected on either. The P3 reading is measured on zero-FTS queries only and does not apply to lexical-anchor queries, whose pages the FTS leg fills; its negative side is a page-max score while its positive side is a query→own-gold cosine, comparable only as "the highest similarity this query can reach in this space". Losing semantic reachability is not losing the answer — the FTS leg can still return it, and full-page recall held at primary hit@5 67.2% across floors 0.197–0.300. Neither reading overrides the 1.9 above, which comes from the blind audit. Annotator and author were the same person |
| Lexical admission was tried and does not close the gap | The FTS leg admits a record on a single 3-character window match, and a round dedicated to fixing exactly that did not pass: 16 arms (window-coverage ratio × unit df ceiling) on a 60,066-row Chinese-competition fixture moved per-page distinctContent on 40 lexical hard negatives from 38/40 over the cap down to 9/40, never to the cap of 2. The 9 that survive all carry anchors that genuinely exist in the corpus, so their window coverage is satisfied by construction. Tightening further spent positives instead: primary hit@5 0.813 → 0.713 and primary_gold absent from the page 11 → 19 of 80 | Full readings in benchmark/reports/fts-round/f4-completion.md; the mechanical verdict is no-safe-arm (f4-selection.json). Production behavior is unchanged — the admission variables never shipped, they existed only in a benchmark injection seam. Content-level evidence, not lexical counting, is the registered next direction |
| Semantic search needs the English form | The semantic leg runs ONLY in the semantic-en-v1 space. If the agent omits semantic_query_en or the guardrail refuses it, the search is keyword-only: no query vector is computed and no stored vector is read, so that request loses semantic reranking as well as semantic recall. The raw-v1 space is never scored, because it has no usable relevance threshold — annotated matches and pure noise overlap in it | semanticEnRate and semanticQueryIssues (including missing) in /health / kiro-mem diagnose show how often that happens |
| Only the top 1,000 semantic candidates keep a rank | The candidate pool is scope-wide — every vector in the searched scope gets scored, and by default no time window narrows that, so no record is invisible to the semantic leg because of its age. What is bounded is retention: candidates are read in chunks, scored immediately, and only the best 1,000 keep a rank (plus every keyword hit, at any score). Returned pages were identical to full scoring on every measured corpus, including at 50,000 records where only 163 of 4,618 above-floor candidates survived — a rank in the thousands cannot reach a page of 10. What does change is observability: a record outside the top 1,000 has no recorded semantic rank | Measured at 50,000 records: full-search p95 158–286ms against a 300ms budget, search-loop memory +225…+324MB against 512MB |
| Install step | Copies the bundled embedding model (~23 MB) into ~/.kiro-mem/models | Model ships in the package — no model download |
| Viewer search is keyword-only | The Viewer never fabricates a semantic_query_en, so its search runs the FTS leg alone: a question phrased entirely differently from the record will not find it here, even though the agent's own search could | Use @kiro-mem/search from a Kiro session for semantic recall; the Viewer labels every result with its match_source so the difference is visible |
| Viewer is loopback-only, single user | It is served by the local Worker on 127.0.0.1 and authenticated by the same local token, handed over in the URL fragment. There is no multi-user model, no remote access and no session beyond the browser tab you opened | Run kiro-mem viewer on the machine that holds the data; a closed tab ends the session |
| Local only | No built-in cross-machine sync | Future: git sync or cloud storage |
License
MIT
