dsh-ab-memory
v0.1.17
Published
--- description: "The dsh-ab-memory package: a cross-session, file-backed agent memory with five tools (remember / recall / forget / load_skill / distill) and an always-on light injection." kind: "package-reference" ---
Maintainers
Readme
description: "The dsh-ab-memory package: a cross-session, file-backed agent memory with five tools (remember / recall / forget / load_skill / distill) and an always-on light injection." kind: "package-reference"
dsh-ab-memory
Summary
One host-only package gives this machine's dsh profiles a cross-session, self-improving
memory. Memory is a local Markdown store that follows Claude Code's "the file is the source
of truth" model: every fact, lesson, skill, and wiki note is a Markdown file under
<memoryDir>/, and a single MEMORY.md index keeps the slugs and one-liners. An optional
Node 22 built-in node:sqlite FTS5 index is a disposable retrieval accelerator — no vector
database, no external service, no network when distillProvider: none.
The package registers the wire names on ctx.tools, contributes one always-on routing
section to ctx.systemPrompt, and reads/writes through whatever ctx.fs the composition
mounts. Three passes — automatic recall, automatic offload of oversized tool results, and
automatic consolidation — run without a tool call, so the store is used even when the model
does not think to reach for it. There is no browser card: this is a host-only (backend) tool,
so it declares no dsh.client half and the Web plugin table has nothing to discover. The
manifest declares dsh.bundle.patch for the host row, so one install — and one patch row —
carries it.
# <DSH_HOME>/profiles/web/cordis.patch.yml
- insert:
- id: memory
name: 'dsh-ab-memory'The deployment must also state the package's required Config guards in an id-only row
(see Mounting) — the schema declares every bound .required() on purpose, so a
hidden default would decide for the deployment how much memory it agreed to carry.
Build, test, and publish
pnpm install # standalone: this package ships its own settings-only pnpm-workspace.yaml
# (it anchors this dir as a standalone workspace root and declares koffi:false
# under allowBuilds), so install stays scoped here and never walks up to the
# harness checkout's root workspace. Do NOT replace node_modules with a symlink
# to another plugin's — that causes "Failed to create bin … ENOENT" WARNs.
pnpm run build # tsc -p tsconfig.json && tsdown -> lib/index.js (node:sqlite stays external)
pnpm test # node --test tests/ (run against the built lib/)
pnpm lint # publint + attw --pack --profile esm-onlyPublishing is the dsh-plugin-release skill's step; its scripts take this package as an
argument. Write <skill> for that skill's base directory, which the loader prints when it
loads the skill:
node <skill>/scripts/src/index.mjs check --package <this package>
node <skill>/scripts/src/index.mjs release --package <this package> --bump patch
node <skill>/scripts/src/index.mjs publish --package <this package> --registry <url>What one call does
Eight tools, one store. remember writes; recall reads; forget tombstones; load_skill
loads a reusable SOP; distill upgrades raw captures into queryable topics; consolidate
merges near-duplicates, downgrades contradictions, and normalizes dates; offload parks a
long text outside the context window; mint crystallizes a stored memory into a skill. Three
of those jobs also happen on their own — see
Automatic recall, offload, and consolidation.
The skill loader is load_skill, not skill: the harness ships its own skill tool for the
session skill catalog, and that registration shadows a same-named global one, which would leave
every stored SOP unreachable.
| Tool | Direction | Step | Detail |
|---|---|---|---|
| remember | Write | Scrub | Runs secret.ts over the content; API keys, inner URLs, and credentials become [redacted] before anything is written |
| remember | Write | Derive | Title from the explicit title or the first line; slug from slugify(title \|\| firstLine) |
| remember | Write | Store | <memoryDir>/<type>/<slug>.md with frontmatter (name, description, type, slug, scope, confidence, modified, ttlDays, tags, deleted, path) plus a MEMORY.md index line |
| remember | Write | Refresh | Recomputes the always-on auto-injection text so the next prompt sees the new fact |
| agent/turn-stopping | Write | Capture | Appends the turn to raw/<date>.md under **user** / **agent** markers, secrets scrubbed and each message clipped |
| agent/turn-stopping | Write | Distill | Extracts atoms from the **user** segments only, skips slugs the store already holds, and refreshes the injection |
| recall | Read | Search | Two-stage: a cheap keyword pass over the index, then an FTS5 pass over the optional sqlite index whenever it exists; returns the top limit (default 5, capped 50) with excerpts and scores |
| recall | Read | Filter | Optional type filter (topic / lesson / skill / wiki); excludes tombstoned records |
| forget | Write | Tombstone | Logical delete (deleted: true + index-line removal); the file is kept for audit because dsh-fs exposes no hard delete |
| forget | Write | Locate | When type is omitted, scans every kind for the slug, then tombstones the one that matches |
| load_skill | Read | Load | Reads a skill asset and returns its full body for the agent to follow |
| distill | Write | Extract | Reads raw/*.md (or one named source), splits free text into atomic assertions with extractAtoms, dedupes against existing topics by slug |
Separating write from read is what keeps the injected context bounded: remember and
distill grow the store, but only the cached, byte-capped index plus the top lessons ever
reach the prompt.
The storage layout
<memoryDir>/MEMORY.md the index: one line per live record (slug, type, one-liner)
<memoryDir>/topics/<slug>.md a fact or distilled atom
<memoryDir>/lessons/<slug>.md a reusable lesson the agent learned
<memoryDir>/skills/<slug>.md a versioned SOP (trigger, steps, verification)
<memoryDir>/wiki/<slug>.md a longer reference note
<memoryDir>/raw/<slug>.md L0 captures awaiting distillation
<memoryDir>/index.sqlite disposable node:sqlite FTS5 accelerator (rebuilt from files)memoryDir may be absolute or relative to the calling session's workspace; resolveMemoryRoot
prefers the session cwd so each project carries its own memory. The index.sqlite is never the
truth — rebuildIndex regenerates it from the Markdown files, and searchMemory falls back to
pure keyword scoring when the DB is absent.
The tools
remember
Persist a fact, preference, correction, or how-to. Arguments: content (required), scope
(project | user | global, default project), type (topic | lesson | skill |
wiki, default topic), title (optional). Returns { stored, slug, path, type }. Sensitive
strings are scrubbed before write; an empty content is rejected.
recall
Recall past memories relevant to the current task. Arguments: query (required), limit
(default 5, capped 50), type (optional filter). Returns { hits: [{ slug, type, path,
excerpt, score }], total, mode }. mode is "fts" when the sqlite index is live, else
"keyword". Use it before re-deriving something the project may already hold.
forget
Forget a memory by slug. Arguments: slug (required), type (optional; when omitted the slug
is searched across kinds). Returns { forgotten, slug }. The record is tombstoned — kept on
disk for audit but excluded from the index and from recall — because dsh-fs exposes no
delete method.
load_skill
Load a versioned reusable skill distilled from past successes. Argument: slug (required).
Returns { slug, body }; throws when no such skill exists. The slug names a skills/ asset in
this plugin's store, not an entry in the session skill catalog.
distill
Distill raw captures (raw/, L0) into atomic topic facts (topics/, L1). Argument: source
(optional file name within raw/, or a path; when omitted, all raw/*.md are distilled).
Returns { extracted, written, skipped }; an atom whose slug already exists is skipped rather
than overwritten. Run it after a session dump lands in raw/.
consolidate
Merge near-duplicate records, downgrade contradictory older ones, rewrite relative dates to
absolute ones, and optionally tombstone expired records. Arguments: dryRun (preview only),
tombstone (also remove expired records). Returns { normalized, merged, conflicted,
tombstonedStale }. The pass is deterministic and offline: it calls no model.
offload
Park a long text outside the context window. Arguments: content (required), task (grouping
label, default task), summary (one-line account; derived from the first line when omitted).
Returns { written, path, pointer }; the pointer is the one line to keep in context, and the
original is read back by path.
mint
Crystallize a stored memory into a reusable skill with a trigger, steps, and a verify rule.
Arguments: slug (required), type (optional). Returns { minted, slug, path, type }. The
source record is tombstoned first, so the index and the FTS table hold exactly one entry, and
re-minting over an existing skill bumps version and keeps the prior body under versions/.
Automatic capture and distillation
The store grows without a tool call. At every turn boundary the plugin appends the turn to
<memoryDir>/raw/<YYYY-MM-DD>.md and then distills it:
- Capture. Both sides of the conversation are written under
**user**/**agent**markers, secrets scrubbed bysecret.tsand each message clipped, so the dump is the L0 evidence a distill reads. Only messages whose source is the person are recorded as the person's; injected context — runtime snapshots, skill catalogs, compaction summaries — is harness machinery, not conversation, and is left out. A session that never carries a human turn is never captured, which keeps subagent sessions out of the store. - Distillation.
distilltakes only the**user**segments of a marked dump. The agent's own prose is re-derivable from the session log, and extracting it would fill the index that every later session injects with text nobody asked to keep. A source without markers is taken whole: it was placed by hand, so its author already chose what belongs there. Topics written on this path carrytags: [distilled, auto], which separates automatic growth from what the model asked to remember. - Which extractor.
distillProvider: ollamasends the segments todistillModel(a non-reasoning localqwen2.5:7bin this deployment, chosen by measurement — see Known Limitations) as one auxiliary model call per queued batch, asked for plain one-line entries and bounded by a ten-minute deadline (the same bound this deployment gives that provider's stream idle timeout). The design's rule is that extraction quality decides every layer above it, so a real model is the intended path. The call's generation policy is stated by the deployment:distillTemperature(0.1 — extraction is a transcription, so the same input should distill to the same entries),distillMaxTokens, anddistillReasoningEffort. Two ends of a call are kept apart rather than collapsed: a call that completes with nothing worth keeping returns an empty list (the model saying nothing), while a call that does not complete normally — capability absent, model not pulled, deadline hit, adapter error, or an adapter that rejects the configured reasoning effort — logs, retries once without the effort when that was the cause, and then parks the model path fordistillCooldownMswhile the deterministic splitter keeps the store growing.distillProvider: noneskips the model entirely and always uses the splitter. - Timing. Capture happens at each turn boundary rather than once at session end, so a session that is killed — or a process that never reaches disposal — has already recorded what it knew, and a topic stays attributable to the turn that produced it. Distillation is idempotent, so re-reading a dump is a no-op for atoms the store already holds. A model-backed distillation is detached: the turn never waits for it, and speech arriving while one batch is being extracted is merged into that root's next batch rather than queued behind it.
Automatic recall, offload, and consolidation
Three passes run without the model calling anything. recall, offload, and consolidate
are tools, and a model reaches for a tool only when it thinks of it; the index is injected on
every request, but the records behind it are not, and nothing parks an oversized tool output.
The plugin decides these three itself and hands the model the result:
- Automatic recall (
autoRecall, on by default) runs inagent/pre-step. It takes the person's own text for the step, extracts query terms, searches the store through the same enginerecalluses (FTS5, keyword fallback), ranks the candidates, and appends one user-role context message carrying the top records with their paths. Ranking segments CJK text into adjacent character bigrams and keeps Latin/digit runs whole, then weights each term by how rare it is across the candidates: a project name or a file name decides, a term every record carries does not. A record is injected only when it matches at least two distinct terms, clearsautoRecallMinScore, and scores at least half the best candidate's score — a shared word alone never injects. One session receives one given set of records once; the same set is not re-offered on the next step. The injected message's source ismemory-auto, so capture never mistakes the plugin's own injection for the person's words and distills it back into the store. - Automatic offload (
autoOffload, on by default) observestool/resultthroughsession/event. A result at or pastautoOffloadThresholdByteshas its full text parked inrefs/tool-offload/<tool>/<slug>.mdand queued; the next step is injected with a notice naming each parked path and its size, and telling the model to read the path instead of re-running the tool. The result itself is already in the log — the loop appended it — so this does not shrink the turn that produced it. What it buys is the next turn and the next session: the output is addressed by file instead of being re-derived. - Automatic consolidation (
autoConsolidateAfterWrites) counts store writes and runs one deterministic consolidate pass after that many writes once the store has been quiet forautoConsolidateIdleMs. The pass needs no model, so it never contends for the extraction provider, and the timer is registered as an effect: it leaves with the plugin fiber.0disables the trigger; theconsolidateCronrow and theconsolidatetool remain available.
Each pass has a switch, so a deployment can keep the model-only behavior: autoRecall: false,
autoOffload: false, autoConsolidateAfterWrites: 0.
The passes are contributed by the plugin when its host process loads it, so a code or config
change reaches new sessions only after that process restarts — dsh web keeps the behavior
it booted with. When automatic recall appears to do nothing at all, check the process first:
this is the usual cause, ahead of any configurability question. Recall also requires the memory
tools to be visible to that agent; a scope the deployment restricted them out of gets no
injection, which is the intended isolation rather than a failure.
Configuration
The store, injection, and extraction fields are required and have no default: a hidden one
would decide for the deployment how much memory it agreed to carry, and the schema enforces it
with .required() so an omitted value fails activation with $.<field>: missing required
value. The automatic-invocation fields are windows with shipped defaults — they decide how
much the plugin shows and when it maintains itself, not what the deployment agreed to carry —
so an existing id-only row keeps booting without them. apply normalizes the config through
the schema first, which keeps the defaults identical whether the Loader validated the row or
the module was mounted directly.
| Field | Group | Meaning |
|---|---|---|
| memoryDir | store | Directory memory is written to; a relative value resolves against the session workspace |
| maxAutoBytes | guard | Hard UTF-8 byte ceiling on the always-on auto-injection text (minimum 256) |
| maxLessons | guard | Most recent lessons surfaced by the auto-injection (minimum 1) |
| staleDays | guard | Days before a record's ttlDays elapses and it is withheld from auto-injection (minimum 1) |
| distillProvider | extraction | none runs the deterministic offline splitter; ollama asks the configured distillModel through the mounted LLM capability |
| distillModel | extraction | The model id distillProvider: ollama calls, e.g. qwen2.5:7b on a local Ollama. A non-reasoning model must also have no reasoningEfforts entry in the provider's model list |
| distillTemperature | extraction | Sampling temperature for extraction (default 0.1): extraction transcribes what the statements already contain, so it wants the most likely wording |
| distillMaxTokens | extraction | Output-token cap for one extraction call (default 1536) |
| distillReasoningEffort | extraction | provider-default leaves the adapter's setting and is required for a model with no reasoning control; low/medium/high asks the adapter for that level (a hybrid-reasoning model otherwise spends its budget thinking) |
| distillCooldownMs | extraction | How long one failed extraction parks the model path (default 300000); 0 retries every batch |
| readScopes | assembly | Which scopes the standing context and automatic recall surface (default project, user, global) |
| consolidateCron | maintenance | Five-field cron for a scheduled consolidate; off (default) mounts no timer |
| autoRecall | automatic | Recall for the person's own words at agent/pre-step and inject it (default true) |
| autoRecallTopK | automatic | Most records one automatic recall injects (default 3) |
| autoRecallMaxTerms | automatic | Most query terms one human message contributes (default 40) |
| autoRecallMinScore | automatic | Share of a query's weight a record must contain (default 0.2; a record is also kept only at half the best score) |
| autoRecallBytes | automatic | Byte ceiling on one automatic-recall block (default 2048) |
| autoOffload | automatic | Park a tool result past the threshold and announce it by path (default true) |
| autoOffloadThresholdBytes | automatic | Result size that triggers parking (default 32768) |
| autoOffloadMaxPerStep | automatic | Most parked outputs one notice announces (default 3) |
| autoConsolidateAfterWrites | automatic | Writes between automatic consolidate passes; 0 disables (default 20) |
| autoConsolidateIdleMs | automatic | Quiet period before that pass runs (default 120000) |
The always-on injection is the only unbounded-risk surface, and maxAutoBytes caps it. The
index plus the top maxLessons lessons are concatenated and then excerpt-clipped to that
ceiling, so the 300s idle timeout is never approached by a swelling memory. Automatic recall
adds one bounded block (autoRecallBytes) and automatic offload one notice per step
(autoOffloadMaxPerStep), each injected at most once per set.
Commands
pnpm build # tsc -p tsconfig.json && tsdown
pnpm test # node --test tests/ (build first)pnpm test runs against lib/, so build first.
Mounting
The package is out-of-tree. Its dsh.bundle.patch (cordis.patch.yml, shipped in the
package) already inserts the host row under id: memory and carries the full config:
block — so the plugin boots and is tunable straight from that one file. To change behavior
(Constraint 6), edit the config: keys in the package's cordis.patch.yml; that is the
single configuration surface. The profile needs at most an id-only config row to override
deployment-specific values — never a second - insert: row. The Config schema declares every
bound .required(), so they are always pinned (the bundle patch ships sensible values); an
id-only row feeds overrides into the plugin the bundle patch already loaded, rather than
registering it twice (which would crash startup with tool "remember" is already registered
/ prompt section "memory:auto" is already registered).
## <package>/dsh-ab-memory/cordis.patch.yml (the configuration surface — edit here)
- insert:
- id: memory
name: 'dsh-ab-memory'
config:
# Storage floor (required): Markdown store location; relative paths resolve
# against the session workspace, so persistence is workspace-scoped.
memoryDir: .dsh-memory
maxAutoBytes: 12288
maxLessons: 16
staleDays: 90
distillProvider: ollama # 'none' = offline deterministic splitter
distillModel: qwen2.5:7b
# Automatic-invocation windows (shipped defaults; override per deployment).
autoRecall: true
autoOffload: true
autoConsolidateAfterWrites: 20## <DSH_HOME>/profiles/web/cordis.patch.yml (optional per-deployment override only)
# id-only config row: the package loads and configures itself via its own bundle patch.
# Add overrides ONLY here — a `- insert:` row with the same id would double-register.
- id: memory
config:
maxAutoBytes: 8192 # example: tighten the always-on injection ceiling// <DSH_HOME>/profiles/web/package.json (dsh.profile.bundles)
"bundles": [ "...", "dsh-ab-memory" ]The bundle name must resolve from the profile's node_modules — resolveBundleDir looks the
package up under the dsh installation and then under <profileDir>/node_modules. For a
published package, declare it as a link: / file: / version dependency so a pnpm
install places it. For local development against the tree-external source at
.dsh/plugins/dsh-ab-memory, the harness resolves the bundle through a directory junction:
# from <DSH_HOME>/profiles/web/node_modules
mklink /J dsh-ab-memory ..\..\..\plugins\dsh-ab-memoryThe bundle row needs no filesystem provider of its own: it reads and writes through whatever
ctx.fs the composition mounts, and respects the session's sandbox fence (see
Writing under the session's sandbox policy).
Writing under the session's sandbox policy
A deployment that mounts @deepseek-ai/dsh-fs-sandbox fences every write by a per-call
sandbox policy, and that policy — not the backend's own default — names the workspace the
calling session runs in. src/sandbox.ts owns that seam. MutationPolicy resolves
ctx.sandboxPolicy for the call and stamps it onto every write through saveText, so a
remembered or distilled file lands inside the session's own workspace; a confining backend
that receives no policy is refused at load (it cannot be caught later because the backend
declares the policy service as its own injection and is never constructed without it).
saveText is the one path every artifact is written through, so a new call site cannot omit
the fence; it restates a refusal with the path, the mode, and the workspace to move under,
keeping the FS_SANDBOX_DENIED code the backend raised.
Model Experience
System-prompt section
memory:auto at order 120, above the built-in tool band and near the top of the system
prompt so it is always visible. It says the agent keeps a long-term memory, that a short
index is already injected, that relevant records are recalled and injected automatically so
recall is for a different query or a kind the injection did not cover, to call remember
when the user confirms a preference or corrects the agent, to call distill after a session
dump lands in raw/, to call offload for a long text, and to call consolidate when the
store has grown. The section is empty in any scope where the memory tools are not visible.
The text below that prose is the session's own workspace: the provider reads the assembly's
agent, resolves that session's memoryDir, and appends the index and top lessons cached for
that root. An assembly with no session gets the routing prose and no workspace's index, so one
project's memories are never injected into another's prompt.
Automatic messages in the request history
Two model-visible messages are not tool results and not the person's words. Both carry
source.kind = 'memory-auto', which is what keeps capture from mistaking them for the person
and distilling them back into the store:
- A recall message (
form: recall) appended to a step whose human text matched stored records: a heading and one- [type] slug — excerpt (path)line per record. - A notice message (
form: notice) on the step after an oversized tool result was parked: one- refs/tool-offload/<tool>/<slug>.md — summary (bytes)line per parked output, with the instruction to read the path rather than re-run the tool.
Tool schema
Eight schemas. remember takes content, scope, type, title. recall takes query,
limit, type, scope. forget takes slug, type. load_skill takes slug. distill
takes source (optional; when omitted, all raw/*.md). consolidate takes dryRun,
tombstone. offload takes content, task, summary. mint takes slug, type.
Tool-call history and result
A remember call followed by one line naming the stored type, slug, and path; a recall
call returning the matched hits with excerpts; a forget call returning whether the slug was
tombstoned; a load_skill call returning the skill body; a distill call returning
extracted → written, skipped; a consolidate call returning the counts it changed; an
offload call returning the parked path and the pointer to keep; a mint call returning
whether a skill was crystallized and where.
Evidence
| File | Proves |
|---|---|
| tests/index.test.mjs | The core pure functions (userSegments included), the secret scrub, writeMemoryWithIndex / forgetMemory / searchMemory, the registered section and tools, the sandbox refusal, per-session injection isolation, turn capture → topic extraction, and the end-to-end remember→recall→forget→load_skill→distill flow through a real mounted Context |
| tests/auto.test.mjs | The automatic passes: CJK bigram terms and Markdown stripping, IDF-weighted ranking, the recall block's budget and paths, RecallGate deduplication, an agent/pre-step recall injected through the real waterfall (and not repeated for the same set), no injection for an unrelated message, autoRecall: false restoring the model-only path, an oversized tool/result parked under refs/tool-offload/<tool>/ and announced on the next step, a short result left alone, and the automatic fields riding their defaults on a six-field row |
| tests/loader-composition.test.mjs | A real cordis.yml boots with the six required fields, the auto-injection section assembles, and a missing/ out-of-range bound refuses registration |
| tests/hmr-safety.test.mjs | The tools and the routing section leave with their contributing fiber |
| tests/conventions.test.mjs | The source rules the skill scaffolds a package with |
| tests/eval-cases.test.mjs | The distillation battery's structure, and that its echo/reference cases store nothing through the deterministic splitter without a model |
| evals/ (see evals/README.md) | Distillation quality against a live model: 18 cases grouped by failure mode, scored through the production prompt and gates, with the splitter's fallback scored beside them |
npm run eval:distill runs the battery against Ollama and prints a scoreboard with the
junk-leak count, the recall pass rate, latency, and tokens; --json writes a report to compare
two parameter sets. evals/README.md carries the case table, the scoring rules, and which knob
each symptom points at.
Known Limitations and Deferred Work
- Retrieval is still keyword, not semantic.
recallruns FTS5 whenindex.sqliteexists and the keyword scorer otherwise;distillProviderselects the extraction engine and does not add an embedding path. A semantic retriever remains deferred. node:sqliteFTS5 needs Node 22+. The package declaresengines.node >= 22.19, andsearchFtsis only reached when the DB is present; on older runtimes the package still loads and degrades to keyword search.index.sqliteis disposable and not the truth. It is rebuilt from the Markdown files viarebuildIndex; a missing or stale DB never loses data, only retrieval speed.forgetis a tombstone, not a deletion.dsh-fsexposes no delete, so the file stays on disk (markeddeleted: true) for audit; nothing physically removes it.- Secrets are scrubbed, not encrypted.
secret.tsredacts API keys, inner URLs, and credential-shaped strings before write; a value the patterns miss is still written in clear text, so do notremembera full secret and expect protection. - Automatic distillation without a model uses the deterministic extractor.
extractAtomssplits sentences on punctuation and a length window, so an automatic topic is sentence-shaped and can be noise; the design's answer is the 27B extraction step, whichdistillProvider: ollamaturns on. Either way every automatic topic carriestags: [distilled, auto], so the growth is auditable and can be filtered or forgotten in bulk. Nothing else caps it: a long session with long prompts writes many topics into the index every later session injects. - Capture reads the session snapshot once per turn.
snapshotEvents()returns the whole log, so capture is O(events) per turn; a session with many thousands of events pays for that each turn. It is a plain array scan, not a log read, and the cursor keeps any of it from being written twice. raw/only grows. Every captured turn is appended and nothing prunes or ages the files out;staleDaysgoverns the injected index, not the dump. It also duplicates part of the durable session log, which is the price of keeping L0 evidence inside the memory store.- The injected text is cached per memory root. It is warmed when the agent is created, so
a session sees its workspace from its first prompt, and refreshed by
remember/forget/distill. A session that outlives a plugin reload keeps its cached text but is not re-warmed (noagent/createdfires again); restartingdsh webis the supported way to pick up a new build. - The index line is the only structured pointer.
searchMemoryreads the index for the keyword pass; a hand-editedMEMORY.mdthat drifts from the files is trusted as-is until the next write refreshes it. - Automatic recall ranks, it does not understand. The score is a term-overlap weight, so a record that shares vocabulary with the question but answers a different one can be injected, and one that answers without sharing vocabulary is missed. The gates (two matched terms, the score floor, half the best score, one set per session, the byte ceiling) bound the damage; a semantic retriever is the deferred fix.
- The extractor is a non-reasoning model because a reasoning one spends the budget thinking.
Measured with
evals/on this deployment's machine (no discrete GPU: Ollama drops the Intel iGPU,/api/psreportssize_vram: 0): the 27B took 387 s cold and 300 s warm for 35 output tokens, and the 9B atdistillReasoningEffort: lowtook 247.8 s per call, filled its wholedistillMaxTokensbudget with reasoning and returned no entry at all — per-turn distillation there extracts nothing. A non-reasoning 7B answers the same case in 3.0 s and stores the fact, which is why this deployment runsqwen2.5:7bwithdistillReasoningEffort: provider-default. A non-reasoning model cannot be sent a reasoning effort at all: the runtime refuses the request before dispatch, so the model entry declares noreasoningEffortsand the config must not name one. Switch toqwen3.6:27bonce a machine can run it, or usedistillProvider: nonefor the deterministic splitter;distillCooldownMsis the backstop that keeps the automatic path from stalling behind whichever model is configured. Runnpm run eval:distillbefore changing any of these numbers. - The configured distill model must actually exist on the provider. The previous value
named a model that was not pulled, and every extraction failed into the deterministic
splitter — which is where the low-value auto-distilled topics came from.
ollama listsettles it; the cooldown log line names the route that failed. - A small local extractor needs the guard rails. Measured here: the 9B path did reach the
model (Ollama's log shows
POST /v1/chat/completionsreturning 200) and still produced an English entry for a Chinese snippet, plus verbatim copies of the input. Two guards answer that — the extraction prompt demands a rewrite in the snippet's language with a worked example, andparseAnswerdrops an entry whose script does not match the snippet's (dominantScript) when both are classifiable. Expect to review the store after changing models: the cooldown and the script gate keep the automatic path from storing the damage but cannot make a weak model good. - A label is not a memory. Measured live after several sessions: 3/3 stored topics were
verbatim echoes of the person's own instructions (
记忆插件提供的工具有没有自动调用,记忆插件发布安装了最新的 服务已经启动 继续) — zero useful entries. Neither the deterministic splitter nor a small model can tell an instruction from a fact, soisStatementLikenow requires clause or sentence punctuation and drops a bare label, title, or pasted command;extractAtomsalso stops splitting on a dot inside a number (0.1.13was being shredded into fake sentences). The cost is stated: a punctuation-free sentence is dropped too — a missing entry costs one lookup, while a stored title costs context in every later session. - The published tarball carries
lib/. A release packed from a tree whose build is stale ships old code, and the resulting behaviour mismatch is easy to blame on config. Build (or runnpm run verify) before packing, and check the installed copy:Select-String <profile>/node_modules/dsh-ab-memory/lib/index.js -Pattern distillCooldownMs. - Automatic offload does not shrink the turn that produced the output. The loop has already appended the tool result, so the full text is in that turn's history regardless. The pass buys the next turn and the next session a path to read instead of a reason to re-run the tool.
- Automatic recall covers the person's own words only. A follow-up instruction that carries no topical text of its own ("继续", "do it") recalls nothing; the records injected for the earlier turn are still in the history, which is the mitigation rather than a new query.
- One automatic recall per record set per session.
RecallGatekeys on the matched set, so re-asking about the same records later in the session relies on the earlier injection still being in the history rather than on a second injection.
Dev Note
pnpm build compiles with tsc and bundles with tsdown; pnpm test runs node --test
over the built lib/index.js, so a green suite means the artifact the profile row resolves is
the artifact that behaves. node:sqlite stays external to the bundle (the package relies on
the Node 22 built-in), so the built entry does not carry a database dependency. The plugin is
host-only: defineTool registrations, the memory:auto section, and the automatic passes
(agent/pre-step, session/event, the debounced consolidate) live in src/index.ts, the
pure core (slug, frontmatter, index lines, scoring, CJK terms, atom extraction) in
src/core.ts, the per-session state the automatic passes need (recall deduplication, the
parked-output queue) in src/auto.ts, the store (read/write/forget/search + sqlite) in
src/store.ts, the sandbox seam in src/sandbox.ts, the secret scrub in src/secret.ts,
and the shared vocabulary in src/types.ts. When you add an asset kind or a retrieval mode,
add it to types.ts, core.ts, and store.ts together, then extend
tests/index.test.mjs; a new automatic pass belongs in src/auto.ts and
tests/auto.test.mjs.
