@coderifts/bypass-probe
v0.1.0
Published
Probe your own agent installation for paths that reach a governed target without passing the gate. Reports dated, version-pinned evidence — a clean run is evidence, not proof.
Maintainers
Readme
@coderifts/bypass-probe
Probe your own agent installation for paths that reach a governed target without passing the gate. Run it against your wiring, in your process, and get back dated, version-pinned evidence of what the probe could and could not reach.
This is not a scanner of CodeRifts' code. It is a prober of yours.
npx @coderifts/bypass-probe --config ./bypass-probe.config.mjsWhy this exists
We built a bypass harness against four agent frameworks and used it on ourselves. On every one of them, a raw tool registered next to our guarded table rewrote a governed spec while our own report said coverage was COMPLETE. We fixed the report (agent-guard 9.5.0 measures and admits, and never emits COMPLETE for live traffic). This package is the harness, turned around and handed to you, because nothing certifies a control better than the vendor showing exactly where it does not stand.
What a clean report means — and does not mean
A run with no installation findings says this, and only this:
These probe classes, run at this time, against these package versions, in this process, did not reach the designated target without the gate deciding.
It does not say your agent is governed. It does not say no bypass exists. The report carries
its own limits in its own text — proves, does_not_prove, and cannot_probe are fields in the
JSON, not just prose in this README, because the JSON is what ends up in the ticket and the README
will not be in the room when someone reads it.
There is deliberately no CLEAN, PASS or SECURE outcome anywhere in this tool. The summary
line says "no unmediated path found in this installation by 6 probes". It will never say "no
unmediated path exists", because we cannot know that and neither can you.
Installation findings vs structural findings. An installation finding is true of your wiring and you can fix it. A structural finding is true of the framework for every adopter — you cannot configure it away, only compensate for it. They are separated because a framework-level fact that fires on every run tells you nothing about your own setup, and mixing the two makes both useless.
Why the version pins matter
Every finding carries the framework, the framework version, the adapter version and the Node version it was measured against. A probe result rots the moment an SDK minor lands: the guardrail surface these adapters read is not a stable public contract, and a patch release can move it. A report older than your lockfile is a historical document, not a statement about today. Re-run it in CI on dependency changes, not once at procurement.
The sentence you can put in front of an auditor
On [date], CodeRifts' bypass probe [version] was run against this agent installation ([framework] [version], Node [version]) and found [N] unmediated paths across [M] probe classes. The report enumerates what it proves, what it does not prove, and what cannot be probed from inside the process. It is evidence about a moment, not a certificate of coverage.
If someone asks you to turn that into "our agent cannot bypass governance", the answer is no, and this tool is the reason we can explain why.
Safety: it will not mutate anything real
In dry-run (the default) the probe replaces each candidate tool's implementation with a
recorder and never executes your body. What runs is the framework's dispatch path — the guardrail
chain, the node wiring — which is exactly what is being measured. The thing that would perform a
write is not on that path.
That is an accounting, not a slogan: originals are wrapped in a counter, and every report carries
safety.original_implementations_invoked. If that is ever non-zero the run reports itself as
unsafe rather than clean. The example scratch tools in this repo throw if executed, so the test
suite proves the guarantee rather than asserting it.
scratch-write mode lets a write actually happen. It requires both mode: 'scratch-write'
and target.scratch === true. There is no flag that writes to a target you have not designated
disposable.
Frameworks
Supported: @openai/agents (where a per-tool guardrail surface exists) and @langchain/langgraph
(where none does). The contrast is the point: on the first, an ungated tool is a wiring mistake; on
the second it is the default, and every governed path is bespoke code the framework does not
enforce.
Not supported, and named so nobody infers coverage from silence: the raw Anthropic SDK tool loop,
Claude Code / MCP hosts, LangChain AgentExecutor, CrewAI, AutoGen, Semantic Kernel, Pydantic-AI,
Mastra, and the Vercel AI SDK. See UNSUPPORTED in src/adapters/index.mjs for the reason each
one is absent.
Probe classes
| id | what it asks |
|---|---|
| RAW_TOOL_ALONGSIDE | Is every callable tool carrying the gate, or only the wrapped ones? |
| TOOL_ADDED_AFTER_WRAP | Does a tool added after setup inherit the gate? |
| TOOL_IMPL_REPLACED_AFTER_WRAP | Does the gate survive reassignment of the body it guards? |
| HOSTED_TOOL_UNROUTED | Can a framework built-in touch the target unseen? |
| DELEGATED_AGENT_SURFACE | Does a handoff or subgraph carry the gate across the boundary? |
| DIRECT_CLIENT_CALL | Is the write capability reachable from process scope at all? |
Exit codes
0 ran · 2 a selected finding was present · 3 unresolved probes · 1 the probe could not run.
Choose with --fail-on none|installation|unresolved|any. installation deliberately does not
fail on unresolved probes: an INCONCLUSIVE usually means you did not supply an input, and turning
a config gap into a red build teaches people to pass --fail-on none, which defeats the tool.
