showreceipts
v0.1.0
Published
Your coding agent said "done". Show receipts: audit agent session logs and reconcile every claim against the tool log — offline, read-only, zero dependencies.
Maintainers
Readme
showreceipts
Your coding agent said "done". Show receipts.
That receipt is the bundled demo scenario (npx showreceipts demo) — synthetic data, not a real session. Yours will look like this, built from your own logs.
Coding agents write a detailed log of everything they do — every command, every
edit, every exit code — and then summarise their own work in prose. The two
don't always agree. showreceipts reads the session logs your agents already
leave on disk and prints, for every session, a receipt: what the agent
claimed in its final message versus what its own tool log proves — the
files it actually changed, the commands it actually ran, the tests it actually
passed (or never ran), what the session cost — plus your personal
false-done rate across sessions.
Offline, deterministic, read-only over the logs, zero runtime dependencies, no API key, no LLM anywhere in the loop. Nothing leaves this machine.
Sixty seconds
npx showreceipts # audit every agent session already on disk
npx showreceipts setup # live receipts at the end of every agent turnThe first command needs no configuration: it finds Claude Code and Codex
sessions under their default directories, scans the last 90 days (--since,
--all to widen) and prints a summary table, the latest receipt and your
false-done rate. The second installs a Stop hook in every harness it finds —
idempotent, with backups, previewable with --dry-run — so a receipt lands in
.showreceipts/last-receipt.md the moment an agent claims it's done.
A note on npx: it adds roughly 0.3–0.6 s of npm overhead on every warm run,
and up to about 1.2 s on a cold npx cache (measured 0.33–1.18 s here).
npm i -g showreceipts removes it, and setup recommends the global install
for hooks.
Upgrade with npm update -g showreceipts, then re-run showreceipts setup so
the hooks' launcher copy updates too. To leave: showreceipts setup --remove
uninstalls the hooks surgically (--remove --all also deletes the launcher;
backups of every config it ever touched are kept under
~/.showreceipts/backups), then npm rm -g showreceipts.
What a receipt looks like
showreceipts demo renders bundled synthetic scenarios so you can see the
output before pointing it at real data. This is the first one (the same
scenario as the SVG above — again, a demo, not a real session):
┌──────────────────────────────────────────────────────────────────────┐
│ RECEIPT #0badf00d · Claude Code 2.1.214 · claude-sonnet-5 │
│ ~/proj/wattage · main · Jul 18 17:14 → 23:52 · 2h 05m │
├──────────────────────────────────────────────────────────────────────┤
│ CLAIMED EVIDENCE │
│ ✗ Lint is clean. ruff check . → exit 1 │
│ (23:44) · never re-run │
│ ? Committed the changes. no git commit in log │
│ (23:52) │
│ ✓ Updated `src/wattage/models.py`. Edit ×3 (17:31, 17:32) │
│ ✓ Created `tests/test_cli.py`. Write (18:05) │
│ ✓ All 41 tests pass. uv run pytest -q → exit 0 │
│ 41 passed (23:41) │
├──────────────────────────────────────────────────────────────────────┤
│ ALSO DID (not mentioned) │
│ · 29 more files changed (src/wattage/, /tmp/demo-scratch/, tests/) │
│ · 3 files written to temp dirs │
├──────────────────────────────────────────────────────────────────────┤
│ 212 tool calls · 31 files changed · 4 test runs · 1 compaction │
│ cost $18.42 (API-equivalent) · cache hit 71% │
│ VERDICT: 1 CONTRADICTED · 1 UNVERIFIED · 3 VERIFIED │
└──────────────────────────────────────────────────────────────────────┘Worst news first: contradicted claims at the top, each with the evidence the
verdict rests on and a timestamp. Below the claims: what the agent did but
never mentioned, the session stats, the cost line and the verdict. All the
demo scenarios — verified, unverified, stale test runs, weakened tests, a
refusal fallback, hook-captured ledgers — are frozen byte-for-byte in
docs/samples/.
Your false-done rate
The audit ends with a table like:
claude-sonnet-5 · claude-code 2.1.214 3/29 done turns contradicted (10%) · 7 unverifiedIt counts done turns, not claims: turns where the agent ended with a final message containing at least one scored claim or a completion marker. A done turn is contradicted when any scored claim in it is contradicted by the tool log, unverified when nothing conclusive was found for at least one claim, clean when every scored claim is verified. Rates are grouped per model × harness × version, and the percentage is hidden below 10 done turns — a small denominator makes a misleading number.
Absence is never contradiction: a claim only gets CONTRADICTED on positive
contrary evidence (a red exit code after the claim, a file the log never
touched). How often the verdicts are right is measured, not asserted — the
hand-labelled precision numbers are in docs/accuracy.md.
What counts as a claim
Claim extraction is a versioned, deterministic rule grammar — no LLM, no
heuristic scoring — applied to the agent's final message only. Negated,
hedged, deferred and quoted-instruction sentences are never scored;
third-party attributions are recognised but not scored. The receipt always
carries the recognized-claim count — claimsRecognized, notScored and
sentencesScanned in the JSON, "claims recognized: N (of which M not
scored)" in the Markdown export, and the terminal no-claims box prints it —
and never implies it scored every sentence.
The full grammar (this table is generated from src/claims/rules.ts and is
the same one the tool executes):
Claim rules (claims/2)
Rules are tried in table order; a clause can yield several claims. Negation, hedge, attribution, scoping and temporal cues (§4.7 step 5) apply to every rule.
| id | kind | trigger | fields | notes |
|---|---|---|---|---|
| test.pass | test | \b(?:all\s+)?(?:the\s+)?(?<!\d/)(?:\d{1,6}\s+)?(?:unit\s+|integration\s+|e2e\s+)?tests?\s+(?:(?:are|is|were|was|still|now|all|already|should|would|might|may|could|will|can|do|does|did|not|never|probably|likely|hopefully|just|no\s+longer)\s+){0,6}(?:pass(?:es|ed|ing)?|green|succeed(?:s|ed)?|ok)(?![\p{L}\p{N}])\btest\s+suite\s+(?:(?:is|was|were)\s+(?:now\s+|already\s+|not\s+)?)?(?:green|passing|clean)(?![\p{L}\p{N}])\btest\s+suite\s+passes\b(?:everything|all)\s+(?:is\s+)?(?:passing|green)(?![\p{L}\p{N}]) | {count?} | Rows 1–2; yields to test.counts when an N/M ratio is in the clause. |
| test.counts | test | (?<![\d/])(?<!\bof\s)(\d{1,6})\s+(?:tests?\s+)?(?:passed|passing)(?![\p{L}\p{N}])\b(\d{1,6})/(\d{1,6})\s+tests?\b(\d{1,6})\s+tests?,\s+0\s+failures?(?:✅|✔|✓|☑|🟢)️?\s*(\d{1,6})\s+passed | {count} · short ratio ⇒ negated {ratio} | Row 2; equal ratio ⇒ positive count, short ratio ⇒ negated. |
| test.gate | test | \b(?:full\s+)?validation\s+gate\b[^.;]{0,140}?\b(?:passed|pass(?:es)?|green|clean)(?![\p{L}\p{N}]) | {count?} + one check {family} per listed tool | Row 1; one check claim per listed tool, count from (N tests. |
| test.count_clean | test | (?<![\d/])\b(\d{1,6})\s+tests?\b(?:\s*([^)]{0,40}))?(?=[^.;]{0,80}\b(?:clean|pass(?:es|ed|ing)?|green|ok|succeed(?:s|ed)?|successful(?:ly)?|no\s+(?:errors|issues|warnings|problems)|0\s+(?:errors|problems)|without\s+(?:errors|warnings)|(?:✅|✔|✓|☑|🟢)️?)(?![\p{L}\p{N}])) | {count} | Rows 1–2; "345 tests, typecheck and lint clean". |
| test.ran | test-ran | \b(?:ran|run|re-?ran|re-?run|running|executed|kicked\s+off)\s+(?:(?:the|all|full|a|any|my|our)\s+){0,2}(?:test\s+suite|tests?|pytest|vitest|jest|specs?|unit\s+tests)(?![\p{L}\p{N}])\btests?\s+to\s+run\b\btest(?:ed)?\s+(?:it|this|that|them|anything|locally|here)(?![\p{L}\p{N}]) | — | Row 3. |
| test.nofail | test | \bno\s+(?:failing|failed|broken|red)\s+tests?\bzero\s+failures?\b0\s+failed\b\bwithout\s+(?:any\s+)?(?:test\s+)?failures? | — | Row 1; consumes its own "no". |
| test.green_marker | test | ^\s*(?:(?:✅|✔|✓|☑|🟢)️?)?\s*(?:all\s+)?green\s*(?:(?:✅|✔|✓|☑|🟢)️?)?\s*[.!]?\s*$ | — | Row 1; "All green ✅" standalone. |
| test.added | test-added | \b(?:added|wrote|created|introduced|implemented)\s+(?:\d{1,6}\s+)?(?:new\s+|more\s+)?(?:unit\s+|integration\s+|e2e\s+|regression\s+)?tests?(?:\s+cases?)?(?![\p{L}\p{N}]) | {count?} | Row 4. |
| check.lint | check | \b(ruff|eslint|flake8|pylint|clippy|golangci-lint|biome|oxlint|rubocop|shellcheck|lint(?:er|ing)?)(?![\p{L}\p{N}])(?:\s+(?:--?[\w-]+|check|run|step|job|output|is|are|was|were|all|and|already|still|now|also|remains?|came\s+back|comes\s+back|→|->|:)){0,4}\s*(?:clean|pass(?:es|ed|ing)?|green|ok|succeed(?:s|ed)?|successful(?:ly)?|no\s+(?:errors|issues|warnings|problems)|0\s+(?:errors|problems)|without\s+(?:errors|warnings)|(?:✅|✔|✓|☑|🟢)️?)(?=\s*(?:[.,;:!)]]|(?:on|for|in|with|across|and|again|at)\b|(|$))\b(?:clean|passing)\s+(ruff|eslint|flake8|pylint|clippy|golangci-lint|biome|oxlint|rubocop|shellcheck|lint(?:er|ing)?)(?![\p{L}\p{N}])\b(lint|ruff|eslint)\s*[::]?\s*(?:(?:✅|✔|✓|☑|🟢)️?|passed|ok)(?![\p{L}\p{N}]) | {family: lint, tool} | Rows 5–6. |
| check.type | check | \b(mypy(?:\s+--strict|[-\s]strict)?|pyright|tsc|typecheck(?:s|ing)?|type-?checks?)(?![\p{L}\p{N}])(?:\s+(?:--?[\w-]+|check|run|step|job|output|is|are|was|were|all|and|already|still|now|also|remains?|came\s+back|comes\s+back|→|->|:)){0,4}\s*(?:clean|pass(?:es|ed|ing)?|green|ok|succeed(?:s|ed)?|successful(?:ly)?|no\s+(?:errors|issues|warnings|problems)|0\s+(?:errors|problems)|without\s+(?:errors|warnings)|(?:✅|✔|✓|☑|🟢)️?)(?=\s*(?:[.,;:!)]]|(?:on|for|in|with|across|and|again|at)\b|(|$))\b(?:clean|passing)\s+(mypy(?:\s+--strict|[-\s]strict)?|pyright|tsc|typecheck(?:s|ing)?|type-?checks?)(?![\p{L}\p{N}]) | {family: type, tool} | Rows 5–6; mypy --strict keeps --strict. |
| check.format | check | \b(prettier|black|isort|ruff\s+format(?:\s+--check)?|gofmt|rustfmt|cargo\s+fmt|biome\s+format|format(?:ter|ting)?)(?![\p{L}\p{N}])(?:\s+(?:--?[\w-]+|check|run|step|job|output|is|are|was|were|all|and|already|still|now|also|remains?|came\s+back|comes\s+back|→|->|:)){0,4}\s*(?:clean|pass(?:es|ed|ing)?|green|ok|succeed(?:s|ed)?|successful(?:ly)?|no\s+(?:errors|issues|warnings|problems)|0\s+(?:errors|problems)|without\s+(?:errors|warnings)|(?:✅|✔|✓|☑|🟢)️?)(?=\s*(?:[.,;:!)]]|(?:on|for|in|with|across|and|again|at)\b|(|$))\b(?:clean|passing)\s+(prettier|black|isort|ruff\s+format(?:\s+--check)?|gofmt|rustfmt|cargo\s+fmt|biome\s+format|format(?:ter|ting)?)(?![\p{L}\p{N}])\b(?:code|files?|everything|tree|it)\s+(?:is|are|was|were)\s+formatted\b\bformatted\s+(?:with|via|using)\s+(\w+)\bformat(?:ting)?\s+(?:check\s+)?(?:is\s+)?(?:clean|pass(?:es|ed|ing)?|green|ok|succeed(?:s|ed)?|successful(?:ly)?|no\s+(?:errors|issues|warnings|problems)|0\s+(?:errors|problems)|without\s+(?:errors|warnings)|(?:✅|✔|✓|☑|🟢)️?)(?![\p{L}\p{N}]) | {family: format, tool} | Row 5. |
| check.build | check | \b((?:npm\s+run\s+|yarn\s+|pnpm\s+|cargo\s+|go\s+|docker\s+|mkdocs\s+|docs\s+|the\s+)build|mkdocs(?:\s+build)?(?:\s+--strict|[-\s]strict)?|webpack|vite\s+build|tsc\s+-b)(?![\p{L}\p{N}])(?:\s+(?:--?[\w-]+|check|run|step|job|output|is|are|was|were|all|and|already|still|now|also|remains?|came\s+back|comes\s+back|→|->|:)){0,4}\s*(?:clean|pass(?:es|ed|ing)?|green|ok|succeed(?:s|ed)?|successful(?:ly)?|no\s+(?:errors|issues|warnings|problems)|0\s+(?:errors|problems)|without\s+(?:errors|warnings)|(?:✅|✔|✓|☑|🟢)️?)(?=\s*(?:[.,;:!)]]|(?:on|for|in|with|across|and|again|at)\b|(|$))^\s*(build)(?![\p{L}\p{N}])(?:\s+(?:--?[\w-]+|check|run|step|job|output|is|are|was|were|all|and|already|still|now|also|remains?|came\s+back|comes\s+back|→|->|:)){0,4}\s*(?:clean|pass(?:es|ed|ing)?|green|ok|succeed(?:s|ed)?|successful(?:ly)?|no\s+(?:errors|issues|warnings|problems)|0\s+(?:errors|problems)|without\s+(?:errors|warnings)|(?:✅|✔|✓|☑|🟢)️?)(?=\s*(?:[.,;:!)]]|(?:on|for|in|with|across|and|again|at)\b|(|$))\b(?:it|code|project|everything|tree|package|crate|module|app|build)\s+compiles\b\bcompiles\s+(?:cleanly|fine|ok|without\s+(?:errors|warnings)|successfully|again)(?![\p{L}\p{N}])\b(mkdocs|docs)\s+build[^.;:]{0,20}(?:clean|pass(?:es|ed|ing)?|green|ok|succeed(?:s|ed)?|successful(?:ly)?|no\s+(?:errors|issues|warnings|problems)|0\s+(?:errors|problems)|without\s+(?:errors|warnings)|(?:✅|✔|✓|☑|🟢)️?) | {family: build, tool} | Row 5; imperative/modal "Build …" is deferred by the cue rules. |
| check.marker | check | \b(lint|typecheck|types|build|format|tests?)\b\s*[:|]?\s*(?:(?:✅|✔|✓|☑|🟢)️?|(?:❌|✘|✗|✖|🔴|❗)️?|passed|ok|clean)(?![\p{L}\p{N}])^\s*(?:(?:✅|✔|✓|☑|🟢)️?|(?:❌|✘|✗|✖|🔴|❗)️?)\s*(lint|typecheck|types|build|format|tests?)(?![\p{L}\p{N}]) | test or check {family} by word | Rows 1/5; "lint ✅, typecheck ✅"; a leading marker with a bare status word. |
| file.verb | file | (?<=^|\b(?:i|i've|we|we've|and|also|then|now|just)\s|[,—:]\s|[,—:])(?<!\b(?:a|an|the|this|that|each|every|any|some|one|newly|previously)\s)(?<!hand-)(?<!\bun)(created|added|wrote|written|generated|scaffolded|introduced|updated|edited|modified|changed|touched|fixed|patched|refactored|rewrote|reworked|cleaned\s+up|removed|deleted|dropped|renamed|moved|extracted|implemented|split)\b\b(?:was|is|has\s+been|have\s+been|were)\s+(created|added|updated|edited|modified|changed|fixed|removed|deleted|renamed|moved|rewritten)\b | {verb, subject, fromPath?, explicitVerb, directObject?} — one claim per PATH | Rows 7–10; one claim per PATH; renamed A to B ⇒ subject B, fromPath A; excluded with "already" or a you/your/the-user subject. |
| file.implemented_in | file | \b(?:is|was|are|were|been|now|i|i've|we|we've|and)\s+(?:implemented|added|defined)\s+in\s+ | {verb: update, subject} | Row 8; UNVERIFIED-only, requires an agent-verb clause. |
| file.count | file-count | \b(\d{1,6})\s+files?\s+(?:changed|modified|updated|touched|edited|created|added)(?![\p{L}\p{N}]) | {count} | Row 11. |
| file.new_file | file | \bnew\s+(?:file|module|test\s+file|component|script|package)\s*[:,]?\s* | {verb: create, subject} | Row 7. |
| command.ran | command | \b(?:ran|run|re-?ran|re-?run|running|executed|invoked|launched|kicked\s+off)\s+\x60([^\x60]{1,120})\x60 | {subject, successPredicate?} | Row 12. |
| command.ran_bare | command | \b(?:ran|run|running|executed)\s+(?:the\s+)?(migrations?|build|linter|formatter|script|smoke\s+test|benchmark|command)(?![\p{L}\p{N}])\b(migrations?|build|linter|formatter|script|smoke\s+test|benchmark)\b[^.;:]{0,30}\bfor\s+you\s+to\s+run\b | {subject} | Row 13. |
| install.pkg | install | \b(installed|added|pulled\s+in)\s+(?:the\s+)?(?:\x60([^\x60\s]+)\x60|((?:@[\w-]+/)?[\w.-]+(?:@[\w.^~-]+)?))\s+(?:as\s+(?:a\s+)?)?(?:dev\s+|peer\s+|optional\s+)?(?:dependency|dependencies|dep|deps|package|packages|plugin|library)(?![\p{L}\p{N}])\b(?:npm|pnpm|yarn|bun|pip|uv|cargo|go|gem|composer|brew)\s+(?:install|add|i)\b[^.;:\x60]{0,40}\x60(\S[^\x60]{0,80})\x60 | {subject} | Row 14; bare nouns only for "added", name@version/@scope for "installed". |
| git.commit | git | (?<!\b(?:a|an|the|this|that|each|every|all|both|your|my|its|their|previously|already|\d+|nine|ten)\s)(?<!hand-)(?<!\bun)\bcommitted\b(?!\s+(?:evidence|artifact|file|fixture|trace|scenario|run|snapshot|recording|version|history|data|baseline|set|transcript|screenshot)s?\b)\bmade\s+(?:a|the|\d{1,6})\s+commits?\b\bcommits?\s+(?:is|are)\s+in\b\bcommit\b(?!\s+(?:message|history|hash|sha|body|trailer|log)s?\b) | {op: commit, sha?} | Row 15; the bare verb needs an agent/negation subject; sha from the clause. |
| git.push | git | \b(?:pushed|push)\b(?=\s*(?:to\b|it\b|up\b|the\s+(?:commit|branch|tag|fix|change)s?\b|\x60|origin\b|(|,|.|—|and\b|$))(?!\s+(?:across|the\s+conversation|back|through|for|on|down)\b)\bis\s+on\s+\x60?origin/[\w./-]+\x60? | {op: push, branch?, remote?} | Row 16; sentence-initial imperative "Push …" is a request, not a claim. |
| git.pr | git | \b(?:opened|created|raised|submitted|filed)\s+(?:a\s+|the\s+)?(?:pull\s+request|PR)\b(?:\s*#?(\d{1,6}))? | {op: pr, prNumber?} | Row 17; a bare "PR #N" is not a claim. |
| git.branch | git | \b(?:created|checked\s+out|switched\s+to)\s+(?:a\s+)?(?:new\s+)?branch\s+\x60?([\w./-]+)\x60? | {op: branch, branch} | Row 18. |
| git.tag | git | (?<![\w-])(?:tagged|created\s+(?:the\s+)?tag|cut\s+(?:a\s+|the\s+)?tag|pushed\s+(?:the\s+)?tag)\s+(?:the\s+)?(?:release\s+|commit\s+|it\s+as\s+)?\x60?(v?\d+(?:.\d+)+[\w.-])\x60? | {op: tag, subject} | Row 18. |
| nochange.marker | no-change | \bno\s+(?:code\s+)?changes?\b(?:\s+(?:to|in)\s+\S{1,60})?\s+(?:were\s+|was\s+)?(?:needed|required|necessary|made)\b\bnothing\s+(?:to\s+change|changed|needed\s+changing)\b\bleft\s+(?:the\s+)?(?:code|files?)\s+(?:as\s+is|untouched|unchanged)\b | {subject?} | Row 21; consumes its own "no"; PATH captured when present. |
| verify.generic | verification | (?:^|\b(?:i|i've|we|we've|and|also|then|now|everything|all|both|each|which|that|it|this|fix\s#?\d+|\w+\s+is|\w+\s+are)\s+)(?:(?:was|were|is|are|has\s+been|have\s+been|got|just|also|then|now|independently)\s+){0,6}(?:re-)?(?:verified|validated|confirmed|double-checked|sanity-checked|smoke-tested)(?![\p{L}\p{N}])(?!\s+(?:email|bug|findings?|context|numbers?|claims?|account|commit|badge|token|user|file|tour|copy|source|data|example|by\s+(?:the\s+)?(?:brief|docs?|user|reviewer|provider|maintainer|community)))\b(?:tested\s+(?:it\s+)?(?:manually|locally|end-to-end|by\s+hand)|manually\s+tested|works?\s+as\s+expected|working\s+(?:correctly|as\s+intended|end-to-end)|confirmed\s+(?:live|working))(?![\p{L}\p{N}])\b(?:should|would|might|may|could|will)\s+(?:now\s+|all\s+|just\s+)?works?(?![\p{L}\p{N}])(?!\s+(?:by|like|around|through)\b) | — | Row 19; never CONTRADICTED; modal "should work" defers via the cue rules. |
| verify.with_cmd | verification | \b(?:verified|confirmed|checked)\s+(?:with|via|using|by\s+running)\s+\x60([^\x60]+)\x60 | verification + command {subject, successPredicate} | Row 20; yields a verification and a command claim. |
| done.marker | completion | ^\s*(?:done|all\s+done|both\s+done|completed?|finished|shipped|that's\s+it)\b(?!\s+(?:when|once|if|for\s+runs))\b(?:it|this|that|everything|all|the\s+\w+(?:\s+\w+){0,3}|phase\s*\w+|step\s*\w+|m\d|fix\s*#?\d+|task\s*\w+|item\s*\w+|milestone\s*\w+)\s+(?:is|are)\s+(?:now\s+)?(?:done|complete|finished|in\s+place)(?![\p{L}\p{N}_])(?!\s*(?:when|once|if|for)\b)\b(?:all\s+)?(?:\d+\s+)?(?:items?|tasks?|steps?|fixes)\s+(?:are\s+)?done\b\bis\s+ready\s+to\s+(?:ship|merge|submit|send|release|publish)\b | — | Row 22; never matches inside backticks; "ready for review" is excluded. |
Worked examples, polarity rules and --explain-claim are in
docs/claims.md.
Which agents are covered
| harness | receipts from | hook events | exit codes | final message | strict nudge |
|---|---|---|---|---|---|
| Claude Code | transcripts on disk (~/.claude/projects) | Stop · SessionStart · PostToolUse · PostToolUseFailure | parsed from toolUseResult | transcript final message | yes |
| Codex CLI | rollouts on disk (~/.codex/sessions) | Stop | parsed from output headers | rollout agent_message | yes |
| Cursor | hook-captured ledger | sessionStart · postToolUse · postToolUseFailure · afterFileEdit · afterMCPExecution · afterAgentResponse · subagentStop · stop · sessionEnd | harness (tool_output.exitCode) | afterAgentResponse.text | yes |
| Gemini CLI | hook-captured ledger | SessionStart · AfterTool · AfterAgent · SessionEnd | parsed (Exit Code: in llmContent) | AfterAgent.prompt_response | experimental |
| Copilot CLI | hook-captured ledger | sessionStart · postToolUse · postToolUseFailure · agentStop · sessionEnd | parsed (exit code N in textResultForLlm) | transcript at agentStop.transcriptPath (best-effort) | no (v1) |
| Hermes | hook-captured ledger | post_tool_call · post_llm_call · on_session_start · on_session_end · on_session_finalize | parsed (extra.status / returncode) | post_llm_call.assistant_response | no (v1) |
| dsh | hook-captured ledger (opt-in) | Stop · SessionStart · PostToolUse · PostToolUseFailure | parsed (Exit code N) | Stop last_assistant_message | as Claude Code (unverified) |
| OpenCode | plugin template (roadmap) | tool.execute.after · session.idle | — | — | no |
| OpenClaw | plugin template (roadmap) | after_tool_call · agent_end · session_start · session_end | — | — | no |
Generated by scripts/gen-docs.mjs from the dialect registry (src/hook/dialects/index.ts); the source columns follow ARCHITECTURE §9. Transcript harnesses are audited from their own files on disk even with no hook installed; ledger harnesses need showreceipts setup first.
Per-harness setup notes (Codex hook trust, Hermes consent, Gemini config
quirks, Cursor exit codes, Copilot final-text limits) are in
docs/harnesses.md.
Privacy
Nothing leaves this machine. There is no network code path in the package —
no http, no fetch, no telemetry, enforced by a lint over the built output
and a test-time network guard — and the readers are strictly read-only over
the harness logs.
What is read: Claude Code transcripts, Codex rollouts, showreceipts' own
hook-captured ledgers, and (for setup/doctor only) the harness config
files. Never read: persisted tool outputs (tool-results/), background task
files (tasks/), auth.json, .env, git internals. What is written: only
.showreceipts/ (under the git root, or the current directory outside a
repo) and ~/.showreceipts/ — receipts, the parse cache, ledgers, backups —
all 0600/0700, atomic and masked.
report --hash-paths replaces every path with a hash so reports can be
shared; bench --publish writes aggregates only (no paths, ids, prompts or
day-precision dates) and never sends anything anywhere. The full tables —
every path read, every file written, every field of the publish payload — are
in docs/privacy.md.
What the cost line means
cost $18.42 (API-equivalent) · cache hit 71% prices the token usage the
harness itself logged against a dated price table
(docs/prices.md, provenance on every row). It is what the
session would have cost at API list prices — subscription users pay their
plan, not this number. It covers only the API calls present in the transcript:
WebFetch/WebSearch sub-requests, compaction and title generation are not
logged by the harnesses and not counted. ≈ marks estimates; unpriced models
are listed, never guessed.
Platforms
macOS and Linux are the supported platforms; Windows is best-effort (the CI job is non-blocking). Node ≥ 20 — Node 20 reached end of life in April 2026, but it stays supported here until showreceipts 1.0.
The rest of the docs
docs/accuracy.md— measured verdict precision and coverage over hand-labelled real sessions, plus known misfires.docs/claims.md— the claim grammar, polarity, worked examples,--explain-claim.docs/harnesses.md— the coverage matrix with per-harness setup notes.docs/privacy.md— everything read, everything written, the publish payload, threat notes.docs/prices.md— the price table with provenance columns.docs/ledger-format.md— the hook-captured ledger format.docs/receipt-schema.md— the--jsonoutput schemas.docs/catalogue.md— the record catalogue generated from the test fixtures.docs/release.md— how releases are built and published.INSTALL_FOR_AGENTS.md— a bootstrap you can paste into a coding agent, with a verify gate.CONTRIBUTING.md,SECURITY.md,CHANGELOG.md.
A 30-second GIF of the demo lives at docs/demo.gif after each release; it is
produced manually with vhs from the
committed demo.tape.
MIT. Source: https://github.com/faizannraza/showreceipts.
