npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@flowstatepm/cortex

v0.1.3

Published

Working memory for AI agents — persistent cognitive state per project

Readme

Cortex

A trust, freshness, and economy layer for coding-agent memory.

Free and open source (MIT). No account, no API key, no telemetry, no paid tier — it runs entirely on your machine and the memory it keeps never leaves it. Install is npm install -g @flowstatepm/cortex plus one setup command; see Install.

Built for Claude Code, and built with it: I wrote it as a solo project using Claude Code as the development environment throughout, and it targets Claude Code as its host — the hooks it installs and the tools it exposes are Claude Code's. Along the way it ran on itself, so most of the design decisions recorded below were ones Cortex was holding for me while I made the next one.

What it does, in two sentences. Claude Code already remembers things between sessions. Cortex is a layer on top that tracks whether each remembered thing is still agreed, still true of your code, and what it cost to inject — so an agent is told "this decision is contested" or "this note points at a file you deleted" instead of being handed a stale fact as settled context.

Where Claude Code actually helped, specifically, since "built with AI" on its own says nothing: the test suite is ~1,800 cases written alongside the code rather than after it; every substantial change went through a review pass where separate agent instances were pointed at the diff with instructions to disprove the previous one's conclusion, which is how several of the defects recorded in docs/invariants.md were caught before shipping. That file is the honest log — measured failures, ideas that were withdrawn, and gaps that are still open.

Agent memory is no longer hard to come by — as of 2026-08-01 Claude Code ships it on by default and other agents ship comparable features. What none of them tell you is whether the thing they just remembered is still true.

Cortex answers three questions a storage layer normally leaves open:

  • Trust — is this still what we decided? Writing a note that opposes an active decision on the same subject marks both sides [contested] on every retrieval surface — recall, brief, state, the SessionStart brief and the reflex — the last two mattering most, since they inject a lone remembered item as settled context. (The operator listing cortex list-memory shows neither marker; cortex inspect-memory reports the contest in full.) Neither side is quietly retired until you close the contest with cortex_resolve. When a decision is genuinely superseded it cools one tier instead of vanishing, and carries a (superseded) label so it can never read as live.
  • Freshness — does this still describe the checkout? File references in memory are validated against the working tree. A memory pointing at a deleted file is penalised in ranking and renders [stale: missing …] rather than disappearing; a renamed file resolves through a git rename map to [moved: a.ts → b.ts]. The penalty is graduated and capped so stale memory stays reachable, which means the label is the guarantee, not the rank.
  • Economy — what did remembering cost? Every channel carries a token budget it actually spends rather than treats as advice: 150 for the session brief, 450 for cortex_brief, 600 for cortex_recall, 800 for cortex_state. Recall and the brief drop evidence from the bottom; cortex_state drops whole sections. Output size is gated retrospectively too — npm run gate fails CI on any rise in output_tokens against a locked baseline, so recall getting more expensive is a build failure rather than something noticed later. cortex stats reports injected/saved/net and a floored saved/injected ratio — for the most recent session and cumulatively for the scope — plus retrieval health (items by state, the count never retrieved, the ten most-retrieved), and books credit only against recorded evidence, so the savings figure is falsifiable rather than modeled.

Under all three: memory is scoped to a branch and a worktree, capture is ambient and costs no process per tool call, and the default lexical ranking is locked against reference baselines that CI re-checks on every push. (The optional CORTEX_SEMANTIC_MODE=rank path is off by default and is not covered by those baselines.)

Saved: now reports zero, and that is the honest number. It used to read 657.6k at 93% efficiency, derived from a single estimate: the difference between a session summary and pasting every captured event into the context as raw JSON. No agent would ever have done that, so the credit was against an action that was never going to be taken — and the quantity was not a context-token saving in the first place, since captured events are internal and never injected. That credit is withdrawn. Credit is now written only with recorded evidence — the file and byte size of a read that genuinely did not happen — and the mechanism that produces it (verified read substitution) has not shipped yet. A falsifiable zero is worth more than an unfalsifiable 93%.

How this compares

Competitors move, so each claim below carries how it was checked. Treat the specifics as dated and the axis of difference as the durable part. Where a cell says not established, it means exactly that: I have not run the thing, and I am not going to assert a limitation I did not observe.

| | Claude Code's own memory | claude-mem | Cortex | |---|---|---|---| | To install | nothing, it is built in | npm | npm, plus jq | | Where memory lives | Markdown files you can edit | SQLite | SQLite | | How a session becomes memory | written as prose | compressed by an LLM | projected deterministically | | Separated per git branch | no — the key is a directory path | not established | yes, branch and worktree | | Flags a memory that contradicts another | no | not established | yes, both sides marked | | Flags a memory pointing at deleted or renamed code | no | not established | yes, [stale: / [moved: | | Reports what remembering cost you in tokens | no | not established | yes, and budgets are enforced | | Retrieval quality checked in CI | n/a | not established | yes, locked baselines | | Editable by hand in a text editor | yes | no | no — via CLI commands |

The honest summary of that table: Claude Code's built-in memory is easier and hand-editable, and if what you want is an agent that recalls your preferences, it is enough. Cortex is for the case where being told how much to trust a memory matters more than the convenience — long-lived branches, decisions that get revisited, codebases where a six-week-old note may describe files that no longer exist.

Native auto-memory writes plain Markdown into a per-project directory and reads it back at session start. It is better than Cortex at three real things: nothing to install, files you can read and edit by hand, and no database to migrate or corrupt. If what you want is an agent that remembers your preferences across sessions, it is enough, and Cortex is not competing for that job.

Its limit is structural rather than a missing feature: the project key is the working-directory path, mangled — ~/.claude/projects/<key>/memory/. (Observed directly, 2026-08-01: one repository's main checkout and its linked worktrees each hold a separate key, and only the main checkout had a memory/ directory at all. That establishes the namespaces are separate; nothing was written in a worktree to test whether main can read it back, so read this as "separate by construction", not as a measured isolation test.) The branch consequence follows from the same fact and needs no measurement: the key is a path, and git switch does not change a path, so a decision recorded while on a feature branch is served back on main as if it had been settled there. There is no branch dimension to have.

claude-mem is the established third-party option and shares Cortex's shape — hooks capture activity, SQLite stores it. The difference is what happens between capture and recall: it compresses sessions with an LLM, where Cortex projects them deterministically. Compression buys better prose about what happened; determinism buys the ability to say why a given answer was returned, to reproduce it, and to gate it in CI. Which you want depends on whether you are reconstructing a session or trusting a decision. It is also the more established of the two by a wide margin, which is worth something concrete: more people have hit its edges before you do. (This paragraph is second-hand — from a 2026-07-28 survey, not from running it. Corrections welcome.)

The six things that are unique here

Each is a behavior you can go and check, not a label:

  1. Branch and worktree scoping. Every worktree of a repository shares one store — the store id is a hash of the absolute realpath of git rev-parse --git-common-dir, and the realpath step is load-bearing rather than tidy: the raw output is .git in a main worktree and an absolute path in a linked one, so only the resolved form is identical across them. Memory inside that store is partitioned by a scope key carrying the worktree path and the branch ref. Switching branches restores the matching snapshot. See Data.
  2. Subagent sessions. A session is identified by (scope_key, agent_id). A subagent gets its own session at dispatch, so even one that reads nothing and runs nothing is attributable; its reads, edits and commands are then filed under that child session, recording parent_session_id and agent_type, instead of being merged into the timeline of the agent that dispatched it — and a tool call carrying an agent_id never rotates or ends its parent's session, whatever its own working directory resolves to. That precondition is load-bearing: a hook installed before agent identity existed sends no agent_id, and those lines take the ordinary primary-session path, which does rotate on a scope change. cortex doctor reports such a hook as out of date. See Subagent attribution.
  3. Deterministic contradiction detection. Contradictions are found by an offline lexical rule — an explicit polarity flip over demonstrably shared context — with no model in the loop, so the same two notes always produce the same verdict. The rules are deliberately strict, because a detector that cries wolf gets ignored. See Contradiction detection.
  4. Checkout freshness with rename resolution. Memory is scored against the current file inventory, with a graduated and capped penalty so stale items stay reachable and labeled instead of buried, and renames resolve to [moved:] rather than counting as missing.
  5. Enforced budgets. The numbers above are spent, not advisory targets a caller is trusted to honour. Two edges, stated because a budget with undisclosed exceptions is only a suggestion: cortex_recall never drops its top-ranked result, so one oversized result overruns the budget rather than returning nothing; and cortex_state skips an over-large section and keeps walking rather than stopping, so a lower-priority section can outlive a higher-priority one. Continuation lines are charged only after every affordable primary line is placed, so adding them cannot change which decisions you see at any budget.
  6. CI-gated retrieval quality. Hermetic seeded suites are locked against reference results; npm run gate fails on a drop in top1_hit or recall_at_3 or a rise in output_tokens, and CI runs it on every push. Suites can also assert on whole rendered surfaces — the SessionStart brief, cortex header, cortex_state — so a change that quietly stops marking a contested decision, or starts leaking a superseded one into the brief, fails the build instead of passing it. Regenerating a baseline requires a Baseline-Regenerated: trailer on the commit that does it. See Retrieval Quality Gate.

How it works

A session, start to finish

When a session starts. Cortex prints a short brief — capped at 150 tokens, and the cap is spent rather than suggested — covering what is load-bearing on this branch: recent decisions, open blockers, what the last session was doing. It also names up to five files this branch has already read that are still unchanged, verified by re-hashing them, so the agent does not spend its opening moves re-reading files to orient itself. On a repository Cortex has never seen, it prints nothing.

While you work. A hook records what happened — files read and edited, commands run and whether they passed, subagents dispatched. This is the part that has to be cheap, because it fires on every tool call: it is pure shell appending a line to a file, with no database opened and no Node process started. Nothing is injected into the conversation here; it is only being written down.

When something gets decided. You (or the agent) record it as a note. If that note opposes an active decision on the same subject, Cortex marks both sides [contested] — it does not pick a winner and does not quietly retire the older one. The contradiction rule is a fixed lexical test with no model involved, so the same two notes always produce the same verdict, and it is deliberately strict: a detector that cries wolf gets ignored.

At the end of a turn. The recorded lines are flushed to the database in one batch. If a subagent ran and reached a conclusion, that conclusion is kept and offered to you as a suggestion to save — it is never written as memory on its own.

When you ask. cortex_recall(topic) searches notes, summaries, snapshots and command history, ranks the results, and renders them inside a token budget. Every line carries its own trust labels: [contested] if it is disputed, (superseded) if it was retired, [stale: missing …] if it points at a file that no longer exists, [moved: a.ts → b.ts] if the file was renamed. Those labels are the product. A memory layer that returns a confident wrong answer is worse than one that returns nothing.

The design in one line

Cortex is retrieval-first and pull-based. It stores decisions, blockers, command outcomes, snapshots, and session summaries in a local SQLite database, captures activity ambiently through hooks, and surfaces remembered content in exactly two channels you did not ask for: a validated session brief at startup, and a short whisper when a high-confidence focus shift matches memory. Two other hooks speak unprompted but inject no memory — a one-line consult hint, at most once per session, and a turn-end nudge when a subagent ran and there are high-confidence notes worth saving. Everything else you ask for.

Core Behavior

  • SessionStart quietly enables capture with cortex inject-header --quiet and prints the validated session brief (or nothing on a cold start). The brief also names up to five files this branch has already read that are still unchanged, most-read first, so a resuming agent does not re-read them to orient itself — verified by re-hashing, and phrased about the files rather than the reader, because a fresh session did not read them.
  • cortex reflect can emit short hook additionalContext on high-confidence focus shifts.
  • Cortex now supports branch-scoped restore: switching branches restores the right snapshot.
  • cortex_route / cortex route provide the cold-callable capability map.
  • Deferred tool discovery should use callable-name discovery (ToolSearch/tool_search query cortex_recall, cortex_state, or cortex_route) or server-name bootstrap (Cortex). Canonical select:mcp__cortex__... selectors may return 0 on current Codex app-server builds and are not proof Cortex is unavailable.
  • cortex_recall(topic) searches notes, summaries, snapshots, and command/episode memory; output is answer-shaped and budgeted (budget, detail: 'scores').
  • cortex_brief(topic) returns a smaller, agent-friendly subset (decisions first, budgeted).
  • cortex_state shows current-session load-bearing notes first, then branch snapshots and the scored working set, within a budget (default 800 tokens).
  • When that state is empty, cortex_state returns fallback guidance instead of an empty string.
  • Note-backed outputs include compact UTC timestamps, for example Decision [2026-06-06 05:18Z]: [auth] use OIDC.
  • Cortex tracks a lightweight current app graph for the active scope and validates file/path references extracted from memory; head changes feed a git rename map.
  • Missing file references demote retrieved memories gently (graduated, capped penalty) and render as [stale: missing ...]; renamed files render as [moved: a.ts → b.ts]; historical queries can still surface them as history.
  • Branch snapshot summaries and recent-session tails prefer notes and file/test/agent activity over raw command-only hook noise.
  • touched and recalled memory stays hot; ignored memory decays out of the default state.
  • resolved notes stay cold and do not trigger hook reflex whispers; cortex_resolve closes them out explicitly.
  • The UserPromptSubmit hook may add a one-line consult hint at most once per session for memory-relevant prompts; calling cortex_route, cortex_recall(topic), cortex_state, cortex_brief, cortex_engage, or topic-based cortex_validate_memory suppresses it.
  • Prompt hooks do not inject memory facts from prompt text; edit and command reflexes still require high-confidence prior context.
  • Optional semantic retrieval is controlled by CORTEX_SEMANTIC_MODE=off|shadow|rank; default is off.

Install

npm install -g @flowstatepm/cortex
cortex install

Mind the scope. Both cortex and cortex-memory on npm are unrelated packages by different authors; neither will give you this tool. The package is @flowstatepm/cortex. The command it installs is cortex — if you already have one of those other packages, its cortex command and this one will collide, so uninstall the other first.

You almost certainly need to install jq first, and it is the one prerequisite that bites. jq is a small command-line JSON reader; the hooks use it because they run on every tool call and starting Node each time would add latency to everything you do. It is not installed by default on any operating system: brew install jq (macOS) · winget install jqlang.jq (Windows) · sudo apt install jq (Debian/Ubuntu).

Without it, cortex install and cortex doctor both fail loudly and name it — but if you skip past that, the hooks cannot parse their input and Cortex records nothing while appearing to work. Measured, not theorised. If a session ever looks like it is remembering nothing, run cortex doctor first; that is what it is for.

Then restart Claude Code, so it picks up the new hook wiring and the MCP server.

cortex install writes the hook scripts with your Node and Cortex paths baked in, merges the wiring into ~/.claude/settings.json, registers the MCP server, adds Cortex's runtime artifacts to the project's .gitignore, and finishes by running the diagnostic. It is idempotent — a second run produces byte-identical files and says Nothing changed. Run it once per project you want Cortex in; the hook wiring is machine-wide, the store is per repository.

On a brand-new project the diagnostic ends with two failures — the store and engagement state, both created by your first session or immediately by cortex inject-header --quiet. That is expected, the report says so, and the command still exits zero, so it is safe in a script.

Installing from a checkout instead

Use this if you want to modify Cortex, or track main directly.

Prerequisites. Cortex's hooks are shell scripts, so two of these are not optional on any platform:

| Requirement | Why | Check | |---|---|---| | Node.js ≥ 18 | the CLI, the MCP server, every hook that does real work | node --version | | git | Cortex scopes memory per repository and per branch | git --version | | bash ≥ 3.2 | every hook script is bash. On Windows, Git Bash — installed with Git for Windows | bash --version | | jq | the hooks parse their JSON payload in shell, before Node | jq --version | | Claude Code | the host that fires the hooks | claude --version |

On Windows, install jq with winget install jqlang.jq. On macOS, brew install jq. On Debian or Ubuntu, sudo apt install jq. Without jq every hook exits silently and Cortex captures nothing — cortex doctor reports it by name.

The bash floor is 3.2 because that is what macOS still ships at /bin/bash, and it is the interpreter the installed wiring names. You do not need a newer one: the hooks use a bash-4 shortcut where the shell has it and an equivalent that works on 3.2 where it does not. Homebrew's bash is fine too, and nothing needs configuring either way.

Then:

git clone https://github.com/ShuromiU/Cortex.git cortex
cd cortex
npm install
npm run build
npm install -g .
cortex install

npm install -g . links the global cortex command to this checkout — it does not copy it. That is deliberate and worth understanding: the checkout is the live installation. npm run build there changes the behaviour of every project on the machine, and switching branches in it changes Cortex everywhere until you rebuild. Installing from npm instead gives you a copy, which does not move when you do.

Verify either route the same way:

cortex doctor

Every failing check names its own fix.

Upgrading an existing machine

cd <your cortex checkout>
git pull
npm install
npm run build
cortex install
cortex doctor

cortex install is what refreshes hook scripts whose templates changed; cortex doctor reports a stale script as a hook-currency failure rather than letting it run silently out of date. Your memory stores live outside the repository at ~/.cortex/projects/, so nothing here touches them, and schema migrations run automatically the first time each store is opened.

Uninstalling

npm uninstall -g cortex-memory

That removes the command and stops the hooks resolving. Your memory is not deleted — the stores stay at ~/.cortex/projects/ until you remove them yourself, and the hook entries stay in ~/.claude/settings.json until you delete them there.

Claude Code Setup

You do not need a CLAUDE.md in every repo just to make Cortex available.

Global Claude settings are enough to:

  • register the MCP server
  • inject Cortex on session start
  • log tool activity through hooks

Use CLAUDE.md only when you want to teach project-specific workflow conventions such as “write blocker notes aggressively” or “brief agents with cortex_brief before delegation.”

MCP Server

Add Cortex to ~/.claude/settings.json:

{
  "mcpServers": {
    "cortex": {
      "command": "cortex",
      "args": ["serve"]
    }
  }
}

SessionStart Hook

Run Cortex quietly at the start of every Claude session:

{
  "hooks": {
    "SessionStart": [
      {
        "matcher": "",
        "hooks": [
          {
            "type": "command",
            "command": "cortex inject-header --quiet"
          }
        ]
      }
    ]
  }
}

cortex inject-header now:

  • consolidates old unconsolidated sessions
  • refreshes branch/project state
  • starts a scoped session
  • auto-engages Cortex for the new session without dumping a large header

Use cortex inject-header without --quiet only when you explicitly want to print the larger branch-aware working-memory header.

Capture, Reflex, Subagent, and Stop Hooks

cortex install writes these for you — see Installing in one command. The wiring it produces:

| Event | Matcher | Script | Cost | |---|---|---|---| | PostToolUse | Read\|Edit\|Write\|Bash\|Agent | cortex-capture.sh | spool append only — no Node spawn | | PreToolUse | Edit\|Write | cortex-reflect.sh reflect-pre | Node only when engaged | | PreToolUse | Agent | cortex-subagent.sh dispatch-pre | one Node spawn per subagent dispatch — never per tool call | | UserPromptSubmit | | cortex-reflect.sh reflect-prompt | Node only when engaged | | SubagentStart | | cortex-subagent.sh subagent-start | one Node spawn per subagent dispatch — never per tool call | | Stop | | cortex-end-of-turn.sh | one Node spawn per turn: spool flush + conditional nudge |

The spool (.cortex.spool.jsonl) is flushed at turn end, at a 256 KiB threshold (detached cortex flush-spool), and at the next session start — leftover lines are never lost.

Subagent attribution

A session is identified by (scope_key, agent_id). The subagent gets its session at dispatch: the SubagentStart hook creates the child before the subagent does anything, so a subagent that only thinks — one that reads nothing and runs nothing — is still attributable. Before this, the only thing that created a child session was captured tool activity, so such a subagent left no trace at all.

From there, each spool line carries the agent_id and agent_type the host reported, so the flush files a subagent's reads, edits and commands under that same child session — recording parent_session_id and inheriting the parent's scope — instead of merging them into the parent's timeline. Both writers find-or-create the same row, so they converge on one session per agent. A line without an agent_id, including every line written by a hook installed before this existed, resolves to the primary session. A subagent's tool call never rotates or ends the session that dispatched it, and ending a session ends its children so they stay reachable by consolidation and GC.

The SubagentStart path creates nothing when the payload carries no agent_id, and nothing when there is no active primary to parent to — silence in both cases, because the alternative is manufacturing a primary session as a side effect of a subagent event. cortex doctor reports a Subagent sessions row once the path has fired at least once, and warns if subagent sessions appear that the hook never saw.

What a subagent is told

A dispatched subagent starts with none of the memory its parent has, and pasting it in by hand is work nobody does. So Cortex briefs it: the dispatch description is recorded when the Agent tool is called, and the matching subagent receives a short brief — at most 150 tokens, the same cap as the session brief — injected into its context before it starts.

Silence is the default, and it is the common case. No matching memory, nothing. A dispatch the hook cannot match to exactly one subagent, nothing. Any failure at all, nothing. Turn the whole thing off with CORTEX_SUBAGENT_BRIEF=off, which stops the recording as well as the brief; cortex_disengage turns it off along with everything else.

Two details worth knowing. The description has to be captured a moment earlier than the subagent starts, because the event that starts a subagent does not carry it — so the wiring above has a second PreToolUse entry, on the Agent tool. And if the parent already pasted the same context into the dispatch prompt, the brief is suppressed rather than sent twice; a parent who paraphrases rather than pastes is not detected, which costs tokens rather than correctness.

When more than one recorded dispatch could be this subagent, nothing is sent. Taking the oldest looks reasonable and is not: a dispatch you deny still gets recorded, and the next same-type subagent in that turn would inherit it — confidently briefed on the wrong job. Refusing costs almost nothing, because the host starts each subagent before recording the next, so an ordinary fan-out never looks ambiguous in the first place. cortex doctor reports how many dispatches were captured, paired and briefed, and how many starts were refused, and warns only if captures accumulate with nothing ever pairing.

The cap is enforced on this brief specifically, rather than inherited from the shared trimming: that trimming always keeps its first item whole, so one long memory would otherwise produce an arbitrarily large brief on a channel nobody asked for.

Two limits worth knowing. Session trees are one level deep: a subagent that itself dispatches a subagent has the top-level session recorded as its parent, not its dispatcher, because "the current session" is primary-only by definition. And the parent is whichever primary is active in the store, which is shared by every worktree of a repository — so with two engaged windows on two worktrees, a dispatch in one can be filed under the other's session. Both are recorded in _bmad-output/implementation-artifacts/deferred-work.md.

What survives a subagent

A dispatched subagent can burn a very large context and leave one paragraph behind. Cortex keeps that paragraph: when the subagent finishes, its final answer is recorded against its own session, and the machinery that decides what is worth remembering can finally see it.

That last part is the whole feature. Cortex decides what a session produced by reading three things — the episodes recorded for it, the events it generated, and the commands it ran — and a subagent's final message was never one of them. For a subagent that mostly thinks and reads, the other two are nearly empty: its file reads record only which lines were opened, and its commands count only when they fail. So the very case subagent sessions exist to make visible produced nothing. Recording the conclusion where that machinery reads is what makes everything downstream work.

It is offered, never written. The conclusion itself is kept automatically — that is a record of what happened, and Cortex records what happens. Anything that looks like a durable decision, blocker or intent is surfaced to you at the end of the turn as a suggestion; it becomes memory only if you choose to save it. A subagent proposes; it does not author. And it is offered once: a suggestion that keeps coming back trains you to dismiss the prompt, which costs more than the prompt ever earned.

Long answers are trimmed at four thousand characters — roughly six hundred words — and the record says when it trimmed, so you are never shown a cut answer that looks complete.

Cortex refuses a subagent retiring memory from an earlier session

Saving a decision retires older decisions on the same subject — that is how memory stays current rather than accumulating contradictions. It also means a subagent, which knows only its own narrow task, could retire something a previous week's work established. Cortex refuses that: a subagent cannot retire, rewrite or delete memory belonging to an earlier session, whether through the Cortex tools or the command line. It gets a plain refusal explaining what to do instead — report the finding, which is recorded automatically, and let the parent act on it.

You are not affected. The refusal applies only to subagents. Your own calls are never blocked, and neither is a subagent acting on memory from the conversation it is part of, or on memory belonging to a different branch. Where Cortex cannot establish with certainty that a target is earlier work on this branch, it allows the call and says nothing: blocking your work by mistake is a worse outcome than the one this prevents.

What it protects, stated exactly: earlier work on this branch that the current conversation did not do. Two limits follow from that and are worth knowing rather than discovering. Cortex's idea of "this conversation" is a bookkeeping record that can outlive a single sitting, so a decision saved a few hours ago under the same record is treated as current and stays reachable. And because the check errs toward letting you through, a subagent determined to route around it can — by assembling a target at runtime, for instance. This is a guard against a subagent doing damage by accident, which is the realistic case; it is not a security boundary, and it does not pretend to be one.

Branch snapshots, scoped session listings and the recent-session tail read primary sessions only; child timelines are reached explicitly. If you upgraded the package but subagent activity still lands on the parent, one of two things is stale: the installed cortex-capture.sh predates the change, or the PostToolUse matcher in your settings no longer lists Agent. cortex doctor reports both — the first as a failing hook-currency check, the second as a capture-matcher warning. Running cortex install fixes both — it rewrites the script and writes the matcher.

Installing in one command

cortex install

Writes the four hook scripts with your Node and Cortex paths baked in, merges the hook wiring into ~/.claude/settings.json, registers the MCP server, and adds Cortex's runtime artifacts to .gitignore — then runs cortex doctor and exits with its verdict. cortex install-hooks is the same command under its old name.

--scope project writes <project>/.claude/settings.json and registers the server in <project>/.mcp.json instead; an unrecognised scope is rejected rather than silently treated as user. --dry-run reports every outcome and writes nothing, and does not run the diagnostic — it has nothing to diagnose. --json emits the result for scripting, with the diagnostic embedded, so a scripted caller can see why a run exited non-zero.

Running it again changes nothing. Not "rewrites the same content" — an unchanged installation produces byte-identical files, so mtimes, backups and your settings formatting are all left alone, and it says Nothing changed.

What it will not do without being told:

  • A hook script you edited is refused, not overwritten. Each installed script carries a digest of the template it came from, and Cortex re-matches the whole template against the file: if it is not exactly what the installer would have written — with any paths — it stops and names --force. With --force, your version is saved to <script>.bak first.
  • A script whose stamp is not this build's is backed up and replaced, not refused. That covers anything installed before stamping existed, and anything from a different version. Refusing would break the repair path cortex doctor recommends. Any overwrite of an existing script keeps a <script>.bak, including the ordinary case where only the baked-in paths changed.
  • An existing cortex MCP registration is left alone. It may point at a checkout you prefer.
  • A settings file that does not parse is refused, never clobbered. Fix the JSON and re-run.
  • A wiring another settings file already provides is not duplicated. Claude Code merges <project>/.claude/settings.json, settings.local.json and ~/.claude/settings.json, so a second entry would not replace the first — both would fire.
  • A hooks directory containing $, a backtick or a backslash is refused. Those are expanded by the shell inside the quoted wiring, so the hook would resolve somewhere else while still looking correct. Spaces are fine.

It does repair what it finds: a PostToolUse matcher that has lost Agent, or a SessionStart command naming a Node that moved, are rewritten in place rather than left alone. Detecting that something is wired is not the same as checking that the wiring is right.

Two things it does change that are worth knowing: your settings file is reformatted to two-space JSON, because it is parsed and re-serialised (comments do not survive — a .bak is written before the first modification), and everything Cortex writes lands via a temp file and a rename, so a half-written settings.json is not possible.

cortex install does not create the memory store or engage Cortex for the project — those happen on your first session, or immediately with cortex inject-header --quiet. Until then the diagnostic reports two failures for that reason and says so.

Diagnosing the installation

cortex doctor

Cortex fails quietly. A hook that never fires, a jq that fell off PATH, a Node that moved — each of them produces an empty memory rather than an error, and an empty memory is indistinguishable from a project you have not worked in yet. cortex doctor is how you tell the difference.

It reports which settings files were readable, engagement state, hook wiring, the PostToolUse capture matcher, hook script presence, placeholder substitution, hook version currency, the configured interpreter, jq, the Node and CLI paths the wiring will invoke, database reachability and schema version, spool size and staleness, and MCP server registration. Every non-passing check names a specific fix — usually a command to run, sometimes an edit to make (a JSON syntax error, a missing binary). --json emits the same report for scripting.

Exit codes: 1 if any check fails, 0 otherwise — so it can gate CI. Warnings do not fail the run; a project you deliberately disengaged with cortex_disengage warns and exits 0.

The check worth knowing about is hook version currency. A hook script installed by an older version stays syntactically valid, correctly substituted, and wired — it simply no longer does what the current build expects. Nothing about it looks broken. cortex install therefore stamps each installed script with a digest of the template it rendered, and doctor recompares that against the template the running build ships. A script with an older stamp, or with no stamp at all, is reported out of date with cortex install as the fix, and fails the run.

Two limits, stated rather than implied:

  • jq and the hook interpreter are located on PATH, not executed. A binary that is present but broken resolves and is reported available. This keeps the run under the 3-second budget without spawning a process per check.
  • Hook scripts are read for their template stamp and their Node invocation, not otherwise validated. A script whose body was edited but which still carries a current stamp and still invokes Cortex is reported healthy. The currency check answers "does this predate the shipped template", not "has this been modified".
  • The diagnostic reads Claude Code's settings files (<project>/.claude/settings.json, <project>/.claude/settings.local.json, ~/.claude/settings.json) and MCP registration from those plus <project>/.mcp.json and ~/.claude.json. Codex wiring is not inspected.

cortex doctor changes nothing it reports on: no session, no engagement write, no schema migration, no spool flush, and no store created where none exists — running it against a project with no database reports the missing database rather than creating one.

It is not literally write-free, and the one exception is worth stating plainly: reading a WAL-mode database creates the .cortex.db-shm and .cortex.db-wal sidecars when they are absent. That is SQLite's requirement for reading WAL at all rather than a choice Cortex makes — a read-only connection prevents content writes, not sidecar creation. The alternative, opening the store as immutable, would read past the WAL and could report a stale schema_version; a wrong answer is worse than a tidy one. On a checkout where those files cannot be created, the database check reports the store as unopenable.

Retrieval Quality Gate

Ranking is benchmarked by hermetic seeded suites in eval/suites/, each locked against a reference result in eval/baselines/. One command runs all of them:

npm run gate

It names the suite and the metric and exits non-zero on a negative top1_hit delta, a negative recall_at_3 delta, or a positive output_tokens delta — accuracy must not fall and output must not get more expensive. noise_count and stale_count are reported for visibility but do not gate.

Everything ambiguous fails closed, because a gate that cannot fail is worse than none:

  • a suite with no baseline, and a baseline with no suite — deleting a suite file must not silently stop gating it
  • a baseline missing a metric — absent values compare as NaN, which would otherwise un-gate that metric while still printing plausible numbers
  • a fixture whose own assertions fail — two suites exist only to lock [stale: and [moved: in the output, and losing a label shrinks the output, so the aggregate delta alone would read the regression as an improvement
  • a suite with no fixtures or no seed, which would score zero on everything and pass forever
  • an unreadable suite, baseline or manifest

Suites can also assert on a whole rendered surface rather than on a recall query, via a surfaces block naming brief, header or full_state with expect_contains / expect_excludes / max_tokens. This exists because those surfaces were previously unreachable: the harness computed header and full_state on every run and asserted on neither, and never built the session brief at all — so the guarantees Cortex publishes about the brief (a contested decision is always marked; a superseded one is never shown) held only by unit test and could regress without a red build. Surface assertions fail on their own terms rather than by a baseline delta, because rendered text is either right or it is not, and a baseline able to record a broken brief as acceptable would defeat the point. An unknown surface name is refused when the suite loads: it would read every assertion against nothing, so contains would fail loudly while excludes and the token budget passed vacuously.

A fixture may also supply its own budget, which is the only way the brief's token-budget enforcement is exercised at all: a seeded brief is well under 150 tokens, so a max_tokens: 150 assertion has too much headroom to ever fire.

One limit worth stating: the brief's read-ledger line is deliberately not gated. It names files that must exist on disk with matching hashes, which a seeded in-memory scenario cannot stage, so leaving it on would make a suite pass or fail by whatever happens to be checked out. That line is covered by unit tests; everything else the brief renders now gates.

It also enforces AD-5: a memory_items kind that no fixture exercises is invisible to the suites rather than penalised by them, so a newly registered kind fails the gate until a fixture ships with it. eval/kind-coverage.json grandfathers the kinds that predate the gate — and a test pins that list to exactly the kinds no suite covers, so widening it means editing an assertion, not quietly appending to an array.

CI runs the gate on every push, after build, lint and tests.

Baselines are locked artifacts. Regenerating one is deliberate:

node dist/transports/cli.js eval-gate --regenerate-baseline budget

The command prints the regressions it is about to bake in, and CI rejects the change unless the commit that makes it carries a Baseline-Regenerated: <reason> line — a trailer elsewhere in the range cannot launder it, and a placeholder reason is rejected. eval/kind-coverage.json is guarded the same way. Regenerating is never the way to turn a red gate green.

MCP Tools

| Tool | Purpose | |------|---------| | cortex_route | Explain ambient memory behavior and route to the right Cortex tool | | cortex_state | Return current-session notes first, then the scored working set, budgeted; empty state returns next-step guidance | | cortex_note | Record an insight, decision, intent, blocker, or focus; reports any active decision the write contradicts | | cortex_resolve | Mark a note resolved or superseded (optionally with replacement content) | | cortex_recall | Retrieve evidence for a topic: lead line + timestamped, validity-labeled evidence within a budget | | cortex_brief | Return a smaller topical brief, optionally for an agent, budgeted | | cortex_suggest_notes | Suggest load-bearing notes from the current session without writing them | | cortex_validate_memory | Audit memories against the current checkout without deleting notes | | cortex_read_ledger | Ask whether files were already read in this scope and whether they changed since — four verdicts, produced by re-hashing | | cortex_search_ledger | Ask whether a search already returned zero results and provably still would — no-matches-at <head>, miss, or unknown | | cortex_engage | Re-enable Cortex if it was disengaged | | cortex_disengage | Disable Cortex hooks for the current session | | cortex_summarize | Force a session summary/checkpoint |

CLI Commands

cortex inject-header
cortex inject-header --quiet
cortex route
cortex reflect --event prompt --prompt "..."
cortex reflect --event edit --file src/file.ts
cortex reflect --event cmd --cmd "npm run test"
cortex status
cortex stats
cortex consolidate
cortex evaluate
cortex evaluate --suite eval/suites/stemming.json --compare eval/baselines/stemming.json
cortex suggest-notes
cortex validate-memory --topic "Activity notes portal"
cortex read-ledger src/db/store.ts src/query/recall.ts
cortex read-ledger src/db/store.ts --json
cortex search-ledger deriveReadKey --path src
cortex search-ledger plainword --glob "*.ts" -i --json
cortex list-memory
cortex list-memory --kind note:decision --state hot,warm --limit 50
cortex list-memory --scope "branch:c:/work/cortex/.git:c:/work/cortex:main" --offset 20 --json
cortex inspect-memory <id>
cortex inspect-memory <id> --json
cortex edit-memory <id> --text "the corrected text"
cortex edit-memory <id> --file correction.txt
cortex delete-memory <id>          # preview; deletes nothing
cortex delete-memory <id> --yes    # actually delete
cortex note-resolve --subject "auth transport" --status superseded
cortex flush-spool
cortex gc            # dry-run report
cortex gc --apply    # actually prune (+ VACUUM when fragmented)
cortex doctor
cortex doctor --json
cortex install
cortex install --dry-run
cortex install --scope project
cortex install --force
cortex serve
cortex log read
cortex log edit
cortex log write
cortex log cmd
cortex log agent

Memory Model

Cortex stores:

  • notes for structured assertions
  • events for raw short-lived activity
  • command_runs for commands plus optional output tails
  • episodes for failures, test cycles, and summaries
  • branch_snapshots and project_snapshots for restore points
  • memory_items as the canonical retrieval/search layer
  • memory_item_semantics for optional summaries, concepts/entities, and JSON-safe embeddings keyed by memory_items.id
  • current_app_graphs for the current checkout's file inventory by scope
  • memory_references for file/path references extracted from memory items

Retrieval is hybrid:

  • FTS over memory_items
  • scope-aware reranking
  • recency/importance/access reinforcement
  • hot/warm/cold decay
  • temporal intent handling for prompts like latest, old, resolved, and when
  • current-checkout reference validation so repo-valid memories beat memories pointing at deleted files or missing plans
  • optional semantic shadow/rank candidates when a semantic provider is configured

Contradiction detection

Writing a note whose content opposes an active decision on the same subject marks both sides contested and returns a payload naming the prior note. The write always succeeds — the conflict is advisory metadata, never a rejection, and choosing a winner stays yours via cortex_resolve.

Detection is deterministic and offline. It fires on an explicit polarity flip — a negation carried by exactly one side, or a curated antonym pair such as enable/disable — and only when the two notes demonstrably talk about the same thing. A detector that cries wolf gets ignored, so misses are the cheaper failure, and the rules are correspondingly strict:

  • A negation must govern something the other note also asserts. "use postgres, not mysql" refines "use postgres" — the negation lands on mysql, which the other note never mentions.
  • Negators are matched on their surface form, never on a stem. Stemming maps noted and noting onto not, which would make "as noted, we cache X" contradict "we cache X".
  • A fragment of a compound is not a negation. --no-verify and src/capture/no-op.ts both contain no.
  • An antonym flip needs near-identical remaining content. "required for rank mode" and "optional for shadow mode" are both true.
  • Overlap is measured against the larger note, so a short one cannot be contained into a match.
  • Detection is scope-keyed. Two branches holding opposite decisions is the ordinary reason branches exist.

Divergent choices (use postgres vs use mysql) and refinements (use postgres vs use postgres with pooling) are not contradictions.

A contested prior is not auto-superseded. Normally a new decision supersedes the old one on that subject, which demotes it one memory tier; suppressing that for contested pairs keeps both sides fully live until you resolve one. That exemption also covers notes already contested, so a later unrelated decision cannot quietly cool one side of an open contest.

Superseded decisions cool instead of vanishing

When a new decision lands on a subject, its predecessor is demoted instead of being archived out of retrieval, as it was before. The durable rule lives in the decay layer: a superseded item always sits one tier below what its activity score would grant — a decision that would derive hot settles at warm, one that would derive warm settles at cold, floored at cold. The old decision stays reachable, demoted in rank below the current one in the ordinary case, and labeled so it cannot read as live:

Decision [2026-07-25 15:56Z]: [queue engine] use kafka for the queue engine.
Decision [2026-06-17 15:56Z]: [queue engine] use rabbitmq for the queue engine. (superseded)

Historical questions — old queue decisions, queue engine history, what did we decide before — reach it through ordinary recall. Blockers are never demoted by a decision: an unresolved blocker on the subject is not superseded guidance.

Because the tier is re-derived from the score on every refresh, recalling a superseded decision cannot reheat it past warm, and it keeps cooling as it ages. Two consequences worth knowing: a freshly superseded decision usually settles at warm (its score is still hot-range), so it can appear in the working set — labeled — until it decays; and the label, not the rank, is the guarantee, since heavy access to the old item can in principle rank it near the new one. The unprompted channels are stricter — the SessionStart brief and the reflex whisper never carry a superseded decision at all, because those channels present a single remembered item as settled context. Manual close-outs behave identically: cortex_resolve(status='superseded') demotes the same way, and pre-existing archived rows from before this behavior stay archived.

Resolving either side with cortex_resolve closes the contest and clears the marker on both. While a contest is open, resolving by subject is refused — picking one of two contested notes would be a guess, and guessing wrong leaves the retracted decision as the live one. Several uncontested notes on a subject (a decision plus a blocker, say) resolve by subject as they always have.

Known limit: token matching is ASCII-only, so non-Latin note content never conflicts. A silent miss rather than a wrong answer.

Contested items in retrieval

A contested memory renders a [contested] marker on every surface that shows it — cortex_recall, cortex_brief, cortex_state (including its Hot: and Resume: lines), the SessionStart brief, and the reflex that injects remembered context automatically:

Decision [2026-07-25 15:56Z]: [spool flush policy] flush the spool at turn end. [contested]
Decision [2026-06-12 15:56Z]: [spool flush policy] do not flush the spool at turn end. [contested]

The unprompted channels matter most. A lone remembered decision injected at session start, or as a reflex, reads as settled fact; if it is one half of an open disagreement, that is precisely the failure the marker exists to prevent.

Both sides are seated together so they read as one disagreement rather than two unrelated claims separated by whatever happened to rank between them. cortex_recall is a flat ranked list and can always do this. cortex_brief and cortex_state sort by note kind first, so they group within a kind — the common case, since a contest always starts from a decision. A contest that spans two kinds stays split there, because seating them together would mean dismantling the kind ordering those surfaces exist to provide.

The marker costs three tokens and is trimmed with its line like any other content. Note that seating both sides together does reorder results: a contested counterpart is pulled up past whatever ranked between the two sides, so under a tight budget it can be kept while a higher-ranked uncontested item is dropped. That is the deliberate resolution of showing a whole disagreement versus showing strictly the best matches.

Rejected alternatives in retrieval

A note written with alternatives renders them beneath the decision they lost to, so an agent about to re-propose one can see it was already considered:

Decision [2026-07-25 15:56Z]: [auth strategy] use OIDC with server-side sessions.
  already rejected: session cookies (no SSO path), JWT-in-localStorage (XSS surface)

This appears in cortex_recall and cortex_brief — the two surfaces where an agent asks a question before proposing an approach. It deliberately does not appear in cortex_state, the SessionStart brief, or the reflex whisper. Those channels budget whole sections or truncate to a fixed width, so an extra line there could not be dropped independently of the decision above it, which is the property the whole feature turns on.

An alternatives line never costs you a decision. Output is assembled in two passes: every decision line that fits is placed first, and only the budget left over buys alternatives. Adding alternatives to a result set therefore cannot change which decisions are rendered, at any budget — the line is charged only once every decision that fits is already on the page. In a two-decision recall the lines cost 47 tokens where there is room and nothing at all once the budget binds.

Two things that follow, and are easy to misread from the output alone. Results are still trimmed from the bottom for their own length, so you can see a decision trimmed while a higher-ranked decision shows its alternatives — the trim was not paid for by that line, and dropping it would not have bought the missing decision back. And because extra budget buys another decision line before it buys alternatives, raising the budget can replace an alternatives line you already had with a further result.

Written as cortex_note(kind='decision', alternatives=['…']). Rationale you put in the strings travels with the rejection; internal whitespace is collapsed to keep each list on one line, and a list long enough to crowd out every other result is truncated.

Inspecting what Cortex holds

Retrieval decides what you see. list-memory and inspect-memory show you everything else.

$ cortex list-memory --limit 2
memory items 1-2 of 3 · newest first (created_at DESC, rowid DESC) · filters: none
notes:208ece98-638a-4c79-b649-ab326abe8ef9  warm  Insight   2026-07-28 04:23Z  insight: the porter tokenizer stems queries too
notes:73e34db9-e13c-4950-a6a4-281e917330e9  warm  Decision  2026-07-28 04:23Z  [auth strategy] decision: do not use OIDC with server-side sessions

next page: cortex list-memory --limit 2 --offset 2

Filter with --kind and --state (comma-separated) and --scope (repeat the flag for more than one). --scope is deliberately not comma-split: scope keys embed the worktree path and the branch ref — they look like branch:c:/work/cortex/.git:c:/work/cortex:main — and git permits commas in branch names, so splitting would shatter a legitimate key.

Two defaults are deliberate and differ from every other surface: no state filter is applied, so archived items are listed too, and no scope filter is applied, so other branches' memory is visible. A listing that quietly omits rows cannot answer "what does Cortex actually hold".

Pages are capped: 20 items by default, 200 at most, however large a --limit you pass. When more remain, the footer prints the next page's command with its filter values quoted, so it runs as printed even when a scope key contains spaces — which, on a path like C:/Claude Code/cortex, it does. The ordering criterion is printed in the header, tiebreaker included, rather than left implicit.

inspect-memory <id> takes a memory-item id or the id of the note behind it — including the counterpart ids that the conflict section prints — and shows the full stored text untruncated, alongside the four things retrieval only ever summarises:

trust:      refs OK

conflict:   contested — an unresolved contradiction on this subject
status:     active
  contested with 73e34db9-e13c-4950-a6a4-281e917330e9 (decision, 2026-07-28 04:23Z)
  already rejected: session cookies (no SSO path), JWT-in-localStorage (XSS surface)

references:
  exists   src/transports/cli.ts

access history:
  count 0, last never
  2026-07-28 04:23Z  auth strategy
  (showing at most 10; cortex gc also prunes the retrieval log — the access count is the durable figure)

The trust label is the same one cortex_recall prints in its lead line. It describes the item's references against its own scope's recorded file inventory, which for another branch may be older than that branch's current checkout — so a cross-scope stale refs means "stale as of what Cortex last recorded there", not necessarily "missing on disk today".

The two halves of access history have different durability and can disagree for two separate reasons, both named in the trailer: the list is capped at the most recent 10, and cortex gc trims the retrieval log independently. access_count is the durable figure.

Stored text is printed verbatim except for terminal control characters, which are stripped: captured stderr can carry ESC and lone CR, and this is the first surface that prints text untruncated. --json stays byte-faithful.

Inspect is the only surface that reads notes.conflict directly. Every other renderer recovers the flag from the projected memory text, because memory_items has no conflict column. Inspect has the id, so it joins to the note itself — and when the column and the projection disagree, it says so rather than silently preferring one. That disagreement is invisible everywhere else.

Both commands are read-only in the sense that matters: neither creates a session, and neither touches access counts, so inspecting memory cannot change the ranking it exists to reveal. --json on either emits the same data as a structure; a missing id exits non-zero in both modes and --json gets a parseable {"error":"not_found"} rather than empty output.

Stored strings are author-supplied, so the renderer treats them as content rather than as its own output: alternatives and subjects are each collapsed onto a single line, and the alternatives payload is capped. Without that, an alternative containing a newline could print what looked like a counterpart line inside the conflict section of a note that has no contest at all.

The read ledger

Cortex records a content digest for every file an agent reads. cortex read-ledger (and the cortex_read_ledger tool) turns that into an answer to the question worth asking before a re-read — have I already got this, and is it still current?

$ cortex read-ledger src/scope/keys.ts src/capture/spool.ts src/db/store.ts docs/gone.md
src/scope/keys.ts: unchanged-since 2026-08-02 18:48Z
src/capture/spool.ts: unchanged-since 2026-08-02 18:48Z (read by subagent general-purpose)
src/db/store.ts: edited-by-you-since 2026-08-02 18:48Z
docs/gone.md: changed-since 2026-08-01 09:02Z (missing)

There are exactly four verdicts — unread, unchanged-since, changed-since, edited-by-you-since — and three properties matter more than the list:

unchanged is never inferred. It is asserted only after re-hashing the file and matching the current bytes against the recorded digest. Cortex does not trust mtime, in either direction: a same-second edit keeps the old timestamp with different content, and a git checkout or a restore rewrites the timestamp without changing a byte. Both cases are tested.

Uncertainty resolves to a miss, never to the convenient answer. A deleted file is changed-since (missing), never unchanged. A file that cannot be hashed — unreadable, now a directory, or past the 2 MiB digest ceiling — is changed-since (unverifiable), and so is a file whose record was oversize and therefore carries no hash to compare. Each of those costs one re-read. None of them can license a skip that turns out to be wrong.

"You already have this" is session-bound, even though change detection is not. A digest recorded by a subagent is a perfectly good fact about the file, and it is reported — but attributed to whoever actually read it, as in the read by subagent general-purpose line above. Cortex will not tell you a file is unchanged since you read it when you never did. A read by your own session or by one of its ancestors is yours; a sibling's or a descendant's is not. A read from an earlier session of the same project is reported as read in an earlier session rather than being named, because every non-subagent session carries the same role label and naming it would say nothing.

edited-by-you-since is checked before the content comparison, and that order matters more than it looks. A digest describes the file as of the moment the capture spool was flushed, not the moment it was read — so the ordinary sequence of reading a file and then editing it records the post-edit bytes. The record then matches what is on disk while your context still holds the old content. Comparing content first would answer unchanged there, which is exactly the wrong skip. If your own edit sits between the record and the question, Cortex says so.

The same flush-time window is the one caveat worth stating plainly: a file changed by something outside Cortex between the read and the flush records the changed bytes and will later read as unchanged. The window is one flush interval, and nothing in the ledger can close it.

Refunding a redundant read

The ledger answers a question you have to remember to ask. Substitution acts on the same evidence without being asked: when a Read returns a file you already read in this session and the bytes on disk still hash to what Cortex recorded, the PostToolUse hook replaces the tool's output with a short line. A four-thousand-token re-read becomes about fifty.

It is off until you turn it on, per project:

cortex substitution on

cortex substitution status reports the current state, and cortex substitution off removes it. Turning it on needs current hooks — run cortex install if cortex doctor reports the hook version as out of date.

What you see in place of the file:

[cortex] substituted: src/db/store.ts is byte-identical to the copy already in this session's
context (verified by sha256 just now). Full content ~4210 tokens. Read it again to get the real text.

That last sentence is load-bearing, not politeness. Reading the same file a second time in one turn is never substituted, so a re-read is always the way back to the real bytes — which is what makes replacing a tool result safe at all.

The conditions are deliberately narrow, and every one of them resolves to no substitution when it cannot be established:

  • The record must prove what your read actually returned, not just what is on disk. Digests are recorded when the capture spool flushes — after the turn — so a read that was followed in its turn by an edit of that file, or by any command (commands rewrite files invisibly: formatters, codegen, git pull), is never certified for refunds. A later clean read re-earns it. This is the guard against the worst failure this feature could have: telling you content is "already in your context" when what you read was different.
  • The file is re-hashed at the moment of the substitution and must match the recorded digest — and its current size must match the recorded size befor