@projectpac/source-agent-history
v0.6.0
Published
Part of PAC: @projectpac/source-agent-history.
Downloads
740
Readme
@projectpac/source-agent-history
The principal's own agent history, as sources this adapter keeps in its own folder: what they work on, derived by asking a model to read everything they have said to their coding agent.
A node set up and started has an identity, a spine, and an executor, and nothing to search over. This is what gives it something -- and it is a source adapter because that is exactly what a source adapter is for: something that creates and maintains the principal's own material out of an outside system the node does not own. The outside system here is a directory on the same machine rather than a daemon or an API, which changes nothing about the seam.
What changed is where the material lands. There is no source registry any more,
so a source is a file in this adapter's own folder --
<dataDir>/agent-history/sources/ -- and two verbs are how anything else
learns one exists: inventory() answers with names and the sentence composed
about each, and stage() copies the named ones into the calling plugin's
folder and answers with relative paths. That is what a flow hands to a box.
The transcripts never become sources
This is the load-bearing decision, not a detail. It used to be a rule about
what a source may be: contents left data management only inside an apply,
into an operation adapter, never into a model prompt -- so a design that
ingested transcripts and then summarized them would have had to break the rule
the whole descriptor/data split existed to keep. The skills flow met the same
wall and removed the offending path rather than widening it.
The rule's mechanism is gone and its reason is not. The transcripts stay where
the harness already keeps them and are re-read once per pass; only what a model
concluded is kept. The bytes are an adapter's input, and an adapter's input is
not a source. What would go wrong is now concrete rather than abstract: a flow
can ask for any source here to be staged into its own folder, so writing
transcripts down would put every secret that ever appeared in a cat .env
one stage() away from flow code.
That is also why this is the one source adapter that injects intelligence:
conversation-notes copies an outside system's records, and this one reasons
over them and keeps what it concluded.
What it keeps, and what it will not open
Nothing is sampled and nothing is truncated. Every prompt, every session title, and every line of the model's own prose is kept in full; a history bigger than one run's context becomes more runs, never less material.
What is left out is material that is not the principal's account of their work, and two exclusions carry the weight:
- Tool output. Claude Code records a tool result as a
type: "user"record whose content is an array oftool_resultblocks, so a filter on the record's type keeps almost all of it -- and with it every secret that ever appeared in acat .env. The filter is on the record's shape instead, and a line carrying the marker is skipped before it is ever parsed. On the machine this was written against that is 18,978 lines rejected by a substring scan, which is why a full pass takes a second and a third rather than a minute. - A subagent's briefing. A session's subagents keep their transcripts in a
directory beneath it, and their turns are marked
isSidechain. Those prompts were written by the agent that spawned the subagent -- "You are an INDEPENDENT VERIFIER..." -- so counting them as the principal's would describe how this harness briefs a subagent rather than how the principal asks for things. The work those sessions did is real, and their prose is kept.
Two things are rewritten on the way through. Secrets are redacted as a backstop -- the shape filter above is the control, since the bytes where credentials actually live are never opened. And text shaped like the intelligence loop's own delimiters is defused, because the loop refuses a context carrying one rather than escaping it: without this, every run on the machine of anyone who has worked on PAC would fail on their own transcripts.
Chunks, and why a second pass is cheap
A chunk exists because one run's context is bounded and a history is not. That is its only reason. Each chunk gets a run describing its slice in the profile's own shape, and one more run merges those -- skipped entirely when the history fits in a single chunk, since that run's answer is already the profile.
Chunks are packed out of whole session files, which is what makes a pass incremental. A file's material never straddles two chunks, so a session that grew dirties exactly one of them and every other partial, already paid for, stands. A pass stats every file and re-reads only what moved, so an unchanged history costs no runs at all -- which matters, because the alternative at every fifteen minutes forever would be the most expensive thing on the node and would buy nothing.
Delivery is at-least-once and a pass can be replaced while its runs are outstanding, so a completion naming any pass but the one in hand is ignored. A chunk that fails is retried on a later pass, up to a bound; the merge runs on the partials there are, and the profile says how much of the history it saw.
What comes out
agent-profile, one document: a summary, the domains, languages, tools and topics the history supports, and how this principal works.agent-project-<name>, one per project the history shows, named after the directory its sessions ran in and never after anything the model wrote. That distinction is load-bearing. Taking the name from the model's prose meant every pass re-grouped or re-capitalized one -- "pikanet" became "Pikanet", "jc-tee-vm" became "jc-tee-vm / jc-box" -- which the removal sweep read as one project vanishing and another appearing: eleven sources deleted and recreated after a pass in which two of two hundred and forty files had changed. An unstable name defeats the point of a source being nameable, so the scan is the authority on which projects exist and the model only supplies the prose. A name it invents is discarded; a project it forgets to mention keeps its source; and a project that leaves the history stops being nameable, because a source a flow could stake that is out of date is worse than none.
Both carry a one-line description, which is the whole of what a peer may ever learn about them.
Running it
Nothing happens on its own. There is no schedule and no work at activation:
pac api POST /flows/agent-history/profileThat answers 202 {profiling, chunks, toReason, projects} once it has read the
history and knows what it will cost, and leaves the runs going. Asking again
while a pass is out is a 409 rather than a second set of runs; asking after one
finished, with nothing changed since, is a 200 {profiling: false} rather than a
bill. {"force": true} in the body drops what the last pass concluded and
reasons over everything again.
pac api GET /flows/agent-history/profileis how far it got: the pass state and every chunk's, and never the profile itself -- that is a source, and no route in this node reads one out.
Why it is not on a schedule. A pass over a real history is ten model runs over three quarters of a megabyte each. Measured against a live node and a real model, the first pass cost $9.84 and each later one about $2.20 -- and the file that changes most reliably is the transcript of the session the principal is in right now, so a fifteen-minute schedule re-profiled every fifteen minutes forever. That is a standing charge for re-reading what the node already knows. A profile is worth asking for; it is not worth subscribing to.
It needs a model, because the profile is something a model wrote: a node whose
pac setup named no executor has nothing to reason with. Everything else is the
ordinary install path, and installing it is the consent -- what the principal
reads before agreeing is this manifest. (There is no approval round trip any
more: installing is one step, and nothing asks. That is the node's current
state rather than its design; see the data core plugin in pac-design.)
What a pass is in the middle of -- which session files it has read, and the
chunks still out -- is kept in the record store, so this waits for
@projectpac/records to be installed before it starts, and GET /plugins
says so (waitingFor: ["records"]) until it is.
Read from ~/.claude by default, and the reader seam is per harness: Codex or
Cursor, or a claude.ai data export, is a reader and a config line rather than a
second plugin. That matters more than it looks, because the plugin id names the
folder every source this writes lives in.
One thing a terminal cannot reach: a node has no way to read the principal's
claude.ai conversations. Everything under ~/.claude is Claude Code, the CLI
exposes no conversation retrieval, and the account's own data export is the only
route out -- which is why an export directory is a reader rather than a request.
Material, for a plugin that reasons over it
The profile is what this adapter concluded. @projectpac/source-intents wants
what it concluded from, and rather than read ~/.claude itself it asks: this
adapter also implements @projectpac/material, so units() lists the session
files and read() hands one over as observations -- through the same readers,
the same filters and the same redaction a profile run is shown, in memory and
never written down. Installing this adapter is the consent to read the
transcripts, and a plugin asking here is bound by that consent rather than
holding one of its own; a node without this adapter has no transcripts to offer
anyone.
It also says when the history moved. watch(listener) fires whenever a session
file under <root>/projects appears or grows, off fs.watch, with a minute's
poll over the listing when the folder is not there yet or cannot be watched --
and it keeps trying to watch, so a folder that appears later is watched from
then on. That is not a schedule and starts no run: a consumer is told a session
grew and decides for itself whether that is worth reading.
