@ishouryasonu/patchpilot
v2.0.0
Published
Verification-first agentic bug fixing system
Maintainers
Readme
PatchPilot
Verification-First Agentic Software Engineering Workflow
PatchPilot is an evidence-first, verification-driven software engineering system designed to investigate software bugs, reproduce failures, implement minimal surgical patches, enforce deterministic patch integrity guards, verify fixes against protected oracle tests, diagnose failures, retry safely, and produce auditable evidence-backed engineering ledgers.
1. Problem & Intended User
Intended User
Software engineering teams, open-source maintainers, and automated CI/CD bug-triage pipelines requiring reliable, self-verifying bug fixes.
Current Bottleneck
One-shot LLM coding assistants frequently generate convincing code that fails on subtle edge cases, introduces regressions into existing test suites, tampers with unit test assertions, or falsely claims success without running tests.
Why It Matters
AI code generation is fast, but unverified code incurs heavy human code-review overhead. PatchPilot uses executable verification gates: a patch cannot be accepted unless the required regression and repository tests pass, patch integrity is verified, and the protected oracle test passes.
2. Solution Architecture & Methodology: Baseline vs. Advanced
+----------------------------------+
| PatchPilot CLI |
+----------------------------------+
|
+---------------------------+---------------------------+
| |
+----------------------+ +-----------------------+
| Baseline Solution | | Advanced Solution |
+----------------------+ +-----------------------+
| |
| Single-Shot Coding Agent | 1. Repository Scout
| - list/read/search | 2. Bug Investigator (.repro test)
| - edit & test | 3. Coding Agent (Fix)
| - stop | 4. Patch Integrity Guard
| | 5. Verification Agent (Diff & Tests)
| | 6. Failure Diagnosis (Max 2 Retries)
| | 7. Final Reviewer & Evidence Ledger
| | 8. Protected Oracle Test Runner
+---------------------------+---------------------------+
|
+--------------------------------------+--------------------------------------+
| | |
+--------------+ +---------------+ +---------------+
| Workspace | | Tool Registry | | LLM Provider |
| Manager | | (Sandboxed) | | Abstraction |
+--------------+ +---------------+ +---------------+
| | |
v v v
Isolated Temp Dir Controlled Runner Zod Schema Parser
(Clean Copy of Repo) (Path check & Allowlist) (Provider Agnostic)Baseline Solution (Single-Shot Agent)
A standard coding agent operating in a simple loop up to a max step limit (15 turns). Uses basic file and test execution tools but lacks specialized investigation, patch integrity checks, independent verification, or diagnostic retry stages.
Advanced Solution (Evidence-First 6-Stage Pipeline + Protected Oracle)
- Repository Scout: Analyzes repository structure, entry points, and dependencies to produce a structured
InvestigationContext. - Bug Investigator: Creates temporary reproduction tests (
.repro/repro.test.js), reproduces the failure, and forms a root-cause hypothesis (InvestigationReport). - Coding Agent: Crafts surgical, minimal patches targeting root cause (
PatchResult). - Patch Integrity Guard: Deterministically audits
git difffor test file deletion, test assertion weakening, protected metadata tampering, secrets (sk-*), blocked commands, and workspace escape attempts (PatchIntegrityReport). - Verification Agent: Independently evaluates
git diffand executes test suites (VerificationReport). - Failure Diagnosis: When verification fails, classifies the breakdown and provides actionable feedback to the Coding Agent (up to 2 retries).
- Final Reviewer & Evidence Ledger: Audits minimal diff size, test results, and bug resolution to render a formal decision while logging an auditable
EvidenceLedger. - Protected Oracle Test Runner: Executes protected oracle tests (
OracleTestRunner) stored outside the workspace with SHA-256 hash checksum verification.
3. Sandboxed Tool Layer & Security Policy
All file and command operations execute inside an isolated temporary directory created by WorkspaceManager.
| Tool Name | Scope & Parameters | Security Controls |
|---|---|---|
| listFiles | listFiles(relativePath?: string) | Excludes node_modules, .git, dist, and blocked file patterns. |
| readFile | readFile(relativePath: string) | Blocks access to .env*, .git/, private keys (*.key, id_rsa), and credentials. |
| searchFiles | searchFiles(query: string, targetPath?: string) | Line-by-line workspace text search. |
| writeFile | writeFile(relativePath: string, content: string) | Enforces strict path sandboxing preventing ../ traversal or host modifications. |
| gitDiff | gitDiff() | Captures git diff inside isolated temporary repository. |
| gitStatus | gitStatus() | Captures status output inside workspace. |
| runCommand | runCommand(commandLine: string) | Controlled execution with working directory scoping, timeouts, allowlists, and secret stripping. |
| runTests | runTests(testPath?: string) | Executes local repository test runner. |
4. Benchmark Suite (12 Controlled Bug Cases, Version 1.0.0)
The benchmark consists of 12 distinct, reproducible software engineering bug cases under benchmarks/. All cases use local, zero-dependency Node.js test runners (node --test), requiring zero external network access. Each case includes a protected oracle test stored outside the workspace in benchmarks/case-XX/oracle/.
| Case ID | Case Name | Bug Category | Description |
|---|---|---|---|
| case-01 | duplicate-operation-retry | duplicate operation | Task queue pushes task ID on retry, resulting in duplicate entries. |
| case-02 | null-undefined-handling | null/undefined handling | Unsafe property access on null nested address payload throws TypeError. |
| case-03 | validation-failure | validation failure | Email validator regex permits invalid untrimmed strings with whitespace. |
| case-04 | pagination-off-by-one | pagination | Off-by-one calculation (page * pageSize) skips page 1 items. |
| case-05 | datetime-boundary-overflow | date/time boundary | Jan 31 billing date calculation overflows into March 2 instead of Feb 29. |
| case-06 | stale-cache-invalidation | stale cache | Updating user email leaves old email key in secondary lookup index. |
| case-07 | retry-idempotency-double-debit | retry/idempotency | Payment processor mutates balance on retry attempts, causing double debit. |
| case-08 | incorrect-sorting-null-prices | sorting/filtering | Product price sorter treats null as $0.00, sorting unpriced items at top. |
| case-09 | auth-token-refresh-grace-period | authentication refresh logic | Auth client fails to trigger refreshToken() within 30s expiry grace window. |
| case-10 | calculation-compounding-discount | calculation/business rule | Discount engine sums promo percentages instead of applying compounding discounts. |
| case-11 | cross-module-config-rename | cross-module regression | Shared config renamed maxConnections to connectionLimit; worker still references old key. |
| case-12 | concurrency-out-of-order-events | concurrency/race-condition | (Harder Case) Async state synchronizer overwrites state when out-of-order stale version events resolve. |
5. Primary & Secondary Metrics
Primary Metric: Verified Fix Rate
$$\text{Verified Fix Rate} = \frac{\text{Cases ACCEPTED (Oracle PASS + Repo PASS + Integrity PASS + Workflow ACCEPT)}}{\text{Total Benchmark Cases}}$$
Secondary Metrics
- False Acceptance Rate: Percentage of candidate patches accepted by workflow checks that fail the protected oracle test.
- Unsafe Patch Rate: Percentage of patches with hard integrity violations.
- Evidence Completeness Score: Average percentage (0-100%) of key engineering claims backed by concrete evidence items.
- Integrity Violation Count: Total number of hard patch integrity violations caught.
- Test Pass Rate: Percentage of cases with passing unit test suites.
- Regression Rate: Percentage of cases where candidate fix broke existing tests.
- Average Retries: Mean diagnostic retry attempts per case.
- Average Tool Steps: Average agent turns required.
- Average Execution Time: Wall-clock execution time per run (milliseconds).
- Estimated Cost: USD cost calculated from token usage.
6. Reproduction Guide & Quick Start
Environment Requirements
- Node.js:
>= 18.0.0 - npm:
>= 9.0.0 - TypeScript:
^5.5.2 - OS: macOS / Linux
Setup Instructions
# 1. Clone repository
git clone https://github.com/Shonu72/patchpilot.git
cd patchpilot
# 2. Install dependencies
npm install
# 3. Environment configuration
cp .env.example .env
# Edit .env to set PATCHPILOT_PROVIDER=openai and PATCHPILOT_MODEL=gpt-4o for live LLM benchmarksExecution Commands
# 1. Run full internal unit and integration test suite (16 suites, 61 tests)
npm test
# 2. Validate all 12 benchmark cases
npm run benchmark:validate
# 3. Run baseline evaluation suite across all 12 cases
PATCHPILOT_PROVIDER=mock npm run evaluate:baseline
# 4. Run advanced evaluation suite across all 12 cases
PATCHPILOT_PROVIDER=mock npm run evaluate:advanced
# 5. Run full benchmark evaluation & generate comparison reports
PATCHPILOT_PROVIDER=mock npm run evaluate:all7. User & Workspace Integration
PatchPilot is published on npm as @ishouryasonu/patchpilot:
- Run via npx:
npx @ishouryasonu/patchpilot evaluate all - Custom Project Run:
npx @ishouryasonu/patchpilot advanced --repo /path/to/your/repo --issue "Bug description" - Global Install:
npm install -g @ishouryasonu/patchpilot
See our detailed guides:
- 👉 User & Workspace Integration Guide
- 🎬 Demo Video Presentation Script
- 📋 Technical Design & Validation Summary
8. Experimental Results
⚠️ Note on Provider Execution:
The results below represent Infrastructure Smoke-Test Results using the Deterministic Mock Provider (PATCHPILOT_PROVIDER=mock). Used for CI infrastructure testing and pipeline verification—not presented as final LLM model intelligence evaluation.
Evaluated across all 12 benchmark cases (Benchmark Version 1.0.0):
| Metric | Baseline Solution | Advanced Solution | Delta | |---|---|---|---| | Verified Fix Rate (Primary) | 0.00% (0/12) | 0.00% (0/12) | 0.00% | | False Acceptance Rate | 0.00% | 0.00% | 0.00% | | Unsafe Patch Rate | 0.00% | 0.00% | 0.00% | | Evidence Completeness Score | 0.00% | 100.00% | +100.00% | | Test Pass Rate | 8.33% | 8.33% | 0.00% | | Regression Rate | 91.67% | 91.67% | 0.00% | | Average Retries | 0.00 | 1.83 | +1.83 | | Average Tool Steps | 1.00 | 10.00 | +9.00 | | Average Execution Time | 201ms | 1401ms | +1200ms | | Estimated Cost (USD) | $0.0000 | $0.0000 | $0.0000 |
Key Takeaways from Infrastructure Testing
- Protected Oracle Gating: When mock providers produce fallback responses without real code edits, protected oracle tests fail (
oracleResult.passed = false), ensuring 0% false acceptance rate. - Audit Evidence Trail: Advanced mode generates 100% complete
EvidenceLedgerJSON artifacts underevidence/advanced/*.json.
9. Main Failure Mode & Engineering Hot Take
Main Failure Mode
Context Scope Fragmentation in Cross-Module Dependencies (case-11)
When an issue spans multiple files (e.g. config.js and workerQueue.js), agents inspecting only a single entry point file produce incomplete fixes.
Engineering Hot Take
"A coding agent that claims its patch works without independent test execution evidence is just guessing with confidence. Real agentic software engineering requires empirical verification gates, protected oracle tests, workspace sandboxing, and auditable evidence ledgers."
10. Trajectory & Evidence Disclosures
Complete step-by-step JSONL trajectories and Evidence Ledgers are persisted in:
trajectories/baseline/*.jsonl&trajectories/advanced/*.jsonlevidence/baseline/*.json&evidence/advanced/*.json
All logs automatically scrub sensitive credentials (OPENAI_API_KEY, GEMINI_API_KEY, passwords, secrets) before writing to disk.
