breach-gate
v2.0.0
Published
API deploy gate. Blocks a deployment only on proven exploitation, with EPSS and CISA KEV prioritisation and published precision numbers.
Maintainers
Readme
Breach Gate
An API deploy gate that only blocks on proof.
Breach Gate scans a running HTTP API and returns one of three answers: SAFE, REVIEW_REQUIRED, or UNSAFE. It blocks a deploy only when it can show you the exact response substring that proves an attack worked.
SECURITY VERDICT: UNSAFE TO DEPLOY
Unauthenticated database access via SQL injection on GET /api/items
Proof: sql-error
> You have an error in your SQL syntax near "1' OR '1'='1" at line 1
Verify it yourself:
curl "http://localhost:3000/api/items?id=1'%20OR%20'1'='1"Why this exists
Most scanners answer "what might be wrong?" and hand you 400 findings. Breach Gate answers "can I ship this?" and hands you a decision.
That only works if the decision is trustworthy. A gate that blocks good deploys gets deleted from the pipeline in a week and never reinstated, so the design rule throughout this codebase is:
Nothing blocks a deploy without positive proof. Scanner identity is not proof. Severity labels are not proof. A CVE in a dependency is not proof. Proof is a substring of a response that no correctly behaving API would return.
Measured accuracy
These numbers are regenerated by CI on every run from test/precision.test.ts against a 14-endpoint corpus with declared ground truth. Half the corpus is deliberately clean, including endpoints built to trip naive detectors: a response body containing the word "error", correctly escaped reflection, a 500 with no stack trace, and a server that sets no security headers at all.
| Metric | Result | |---|---| | Precision | 100% (0 false positives) | | Recall | 100% | | Corpus | 6 vulnerable, 8 clean |
Precision is gated at exactly 1.0 in CI. A single false positive fails the build. Recall is gated at 0.9, because a missed finding is a bug but a false positive is an uninstall.
Run it yourself: npm run precision
Scope
Breach Gate does one thing: OpenAPI-aware security testing of running REST APIs.
It deliberately does not scan frontends, GraphQL, or container images. Those were removed in v2.0.0. Six scanners at 40% precision is worth nothing; one scanner you can trust is worth keeping switched on.
| Layer | Tool | Can it confirm an exploit? | |---|---|---| | Dependencies and IaC | Trivy | No. Reports what is present, prioritised by EPSS and KEV. | | Active HTTP probing | OWASP ZAP | No. Reports what it noticed. | | Behavioural testing | LLM-generated tests | Yes, when a response carries proof. |
What counts as proof
A finding is confirmed only when a scanner observed one of these in a response and the same signal was absent from a benign baseline request to the same endpoint:
| Proof | What was observed |
|---|---|
| sql-error | Database error text produced by an injected payload |
| command-output | Output of an injected shell command |
| cloud-metadata | Instance metadata reached via SSRF |
| path-disclosure | File contents or system paths returned |
| payload-reflected | Script payload echoed back unescaped |
| privileged-field | Privileged field accepted via mass assignment |
| auth-bypass | 2xx where the unauthenticated baseline was 4xx |
| timing-oracle | Reproducible response delay |
| stack-trace | Unhandled exception detail returned to the caller |
Every proof carries an excerpt. If Breach Gate cannot quote the evidence, it does not claim the exploit.
Prioritisation
Exploitability comes from real-world data, in this order:
- CISA KEV — the CVE is being exploited in the wild right now
- Proof — we exploited it ourselves during this scan
- EPSS — modelled probability of exploitation in the next 30 days
- Category baseline — a guess, labelled as one in the output
Both feeds are cached locally (.breach-gate-cache, 24h TTL) and fetched once per scan. If they are unreachable the scan continues without them and says so. Use --offline to skip network lookups entirely and rely on the cache.
Scoring
feasibility = reachability × exploitability × impact × confidenceNo floors, no bonuses, no per-source overrides. If a score looks wrong, one of the four factors is wrong and gets fixed there. There is exactly one scoring implementation, in AttackAnalyzer, and reports and verdicts both read from it.
reachability is a path and auth heuristic, not call-graph reachability analysis. It is named that way in the code and it will not be described as more than that.
Verdicts
| Verdict | Exit | Meaning |
|---|---|---|
| SAFE | 0 | Nothing proven, nothing above the review threshold |
| REVIEW_REQUIRED | 0 | Plausible but unproven. Surfaced, does not block. |
| UNSAFE | 1 | Exploitation proven. Deployment blocked. |
| INCONCLUSIVE | 1 | Scan did not complete. Cannot verify, so fails safely. |
A failed scan is never a passing scan. If the target is unreachable, if more than half the test requests error, or if a required scanner fails, the verdict is INCONCLUSIVE and the exit code is non-zero.
Quick start
npm install
npm run build
# Start the deliberately vulnerable demo API
npm run demo
# Scan it
npm run scan -- -t http://localhost:3000 -vUsage
breach-gate scan [options]| Option | Description |
|---|---|
| -c, --config <path> | Config file (default security.config.yml) |
| --configs <paths> | Comma-separated configs for monorepo scans |
| --workdir <path> | Working directory for resolving relative paths |
| -t, --target <url> | Target URL, overrides config |
| -o, --output <dir> | Report output directory |
| -f, --format <formats> | markdown, json, sarif, html |
| --profile <name> | Policy profile: pull-request, main, release, nightly |
| --baseline <path> | Baseline file for tracked waivers |
| --differential | Fail only on findings not in the baseline |
| --ci | Deterministic, minimal output for pipelines |
| --offline | Use cached exploit intel only, no network lookups |
| --explain-verdict | Show every scoring factor and where it came from |
| --skip-static | Skip dependency analysis |
| --skip-dynamic | Skip ZAP |
| --skip-ai | Skip behavioural testing |
| -v, --verbose | Verbose output |
Other commands: breach-gate init, breach-gate doctor, breach-gate watch.
Explaining a verdict
--explain-verdict prints the four factors and the basis for exploitability, so you can see whether a score came from KEV, from a proof, from EPSS, or from a category guess:
Score Finding Factors
0.68 SQLi in items query reach=100% exploit=95% impact=95% conf=90% [CONFIRMED]
└─ Demonstrated during this scan: sql-error
0.31 CVE-2024-21538 in cross-spawn reach= 40% exploit=43% impact=50% conf=100%
└─ EPSS 0.0421 (92nd percentile) probability of exploitation in 30 days
0.02 X-Frame-Options not set reach= 90% exploit=10% impact=15% conf=75%
└─ No exploit intelligence and no demonstration. Category baseline.CI integration
- name: Run Breach Gate
uses: epten08/breachgate@v2
with:
config: security.config.yml
target: ${{ vars.STAGING_API_URL }}
output: security-reports
format: json,sarif
scan-args: --profile main
- uses: github/codeql-action/upload-sarif@v3
if: always()
with:
sarif_file: security-reports/security-report.sarifMore: GitHub Actions, GitLab CI, Azure Pipelines, policy profiles.
AI provider setup
The behavioural scanner needs a model to generate endpoint-aware test cases. The model never decides whether something is a vulnerability; it only proposes what to try.
scanners:
ai:
enabled: true
provider: anthropic # or openai, or ollama for local
model: claude-haiku-4-5-20251001
maxTests: 15# .env
ANTHROPIC_API_KEY=sk-ant-...
# or OPENAI_API_KEY=sk-...
# or run Ollama locally: ollama serve && ollama pull llama3For reproducible pipelines, generate test cases once and replay them:
breach-gate scan --ci # writes security-reports/ai-tests.json
# then set scanners.ai.replayTests to that pathReplay needs no model and no API key, which makes CI deterministic and free.
Suppression
Two mechanisms, for two different situations:
| Mechanism | Use for |
|---|---|
| .breachgateignore | Permanently acceptable: intentional behaviour, headers handled at the CDN |
| .breach-gate-baseline.yml | Temporary waivers with a ticket and an expiry date |
# .breachgateignore
suppressions:
- pattern: "Missing Security Header"
reason: "Set at the CDN, not the origin"
- pattern: "Broken Access Control"
endpoint: "/api/legacy"
reason: "Tracked in SEC-456"
expires: "2026-12-01"You should rarely need these for header noise. Missing headers are reported once, at LOW, with no proof attached, and cannot influence the verdict.
Prerequisites
Node.js 18+. Everything else is optional and degrades gracefully:
| Tool | Needed for | Without it |
|---|---|---|
| Docker | Running Trivy and ZAP as containers | Install them natively, or skip those scanners |
| Trivy | Dependency and IaC analysis | --skip-static |
| OWASP ZAP | Active HTTP probing | --skip-dynamic |
| An LLM key | Behavioural test generation | --skip-ai, or use replayTests |
Run breach-gate doctor to check what is available.
Development
npm run typecheck
npm test # integration
npm run test:controls # negative controls
npm run precision # accuracy measurement
npm run test:all # everything, as CI runs itContributing a detector
Every detector needs a negative control before it can be merged. If you cannot write the test that proves your detector stays quiet on a clean target, you do not understand it well enough to ship it.
The v1 test suite only ever scanned a deliberately vulnerable demo API. UNSAFE was always the expected answer, and UNSAFE is also this tool's failure mode, so the suite could not distinguish working from broken. That is how a missing X-Frame-Options header shipped as a confirmed SQL injection. See test/negative-controls.test.ts.
Migrating from 1.x
- Frontend, GraphQL, and container scanners are removed. Delete those blocks from your config and any
--frontendflags. Finding.riskScoreandFinding.exploitabilityare gone. Scores are computed on demand; JSON and SARIF reports now carryfeasibilityScore,confirmed,proofs,epssScore, andknownExploited.- Far fewer things return
UNSAFE, because unproven findings no longer block. Expect pipelines that were previously red to go green. CheckREVIEW_REQUIREDoutput to see what moved. - ZAP informational and low-confidence alerts are dropped entirely.
License
MIT
