# MCPatch Phase 5.1 benchmark summary

- Benchmark version: `phase5-1.v1`
- Frozen specification checksum: `f716e1bb453857b7bc34f55f05dac6e0a827c163ecf51f015c03c8896081d93b`
- Frozen input commit: `7a747129cbbb954190d0a6e841892ac0f9313535`
- Specification tag: `phase-5-1-benchmark-spec-frozen`
- Randomization seed: `mcpatch-phase5-1-interleave-v1`
- Repetitions: primary 3; calibration 1
- Primary scenarios: 3
- Primary raw trials: 54
- Calibration raw trials (excluded from headline results): 0
- Supported scope: synthetic Python repositories represented by the frozen suite.

## Frozen execution configuration

- Provider: `codex` (`codex-cli 0.137.0`)
- Model: `not observable (the configured Codex model was not exposed)`
- Reasoning setting: `not observable (the configured Codex reasoning setting was not exposed)`
- Runner image: `sha256:ee148fe402624e25323c85366689261b547320d990bc9d9c53e9b9cbb908595b`
- Maximum attempts: 3
- Timeout: 600 seconds
- Resource limits: `{"cpu_count":1.0,"disk_limit_bytes":null,"disk_limit_enforced":false,"memory_bytes":1073741824,"network_policy":"bridge_network_local_datahub_denied","network_policy_enforced":false,"process_limit":128}`

### Conditions and budgets

| Condition | Max attempts | Timeout seconds | Resource policy checksum |
|---|---:|---:|---|
| agent_only_end_to_end | 3 | 600 | `9d04d5137e3145fce1c3939b3834fb6d207bb1c41039efb61f18658262d9d0fb` |
| code_search_discovery | 3 | 600 | `9d04d5137e3145fce1c3939b3834fb6d207bb1c41039efb61f18658262d9d0fb` |
| mcpatch_discovery | 3 | 600 | `9d04d5137e3145fce1c3939b3834fb6d207bb1c41039efb61f18658262d9d0fb` |
| mcpatch_end_to_end | 3 | 600 | `9d04d5137e3145fce1c3939b3834fb6d207bb1c41039efb61f18658262d9d0fb` |
| mcpatch_guided_repair | 3 | 600 | `9d04d5137e3145fce1c3939b3834fb6d207bb1c41039efb61f18658262d9d0fb` |
| repo_only_repair | 3 | 600 | `9d04d5137e3145fce1c3939b3834fb6d207bb1c41039efb61f18658262d9d0fb` |

## Scenario inventory

Calibration scenarios existed in validated Phase 1-4 evidence and are excluded from primary generalization claims.

| Scenario | Tier | Repositories | Frozen workspace layouts |
|---|---|---:|---|
| billing_summary_v2.v1 | primary | 6 | `billing-audit-control`, `billing-console-agent`, `billing-context-provider`, `billing-forecast-control`, `billing-insights-skill`, `collections-assistant-agent` |
| customer_lookup_v2.v1 | calibration | 6 | `customer-insights-skill`, `customer-tool-provider`, `fraud-agent`, `inventory-agent`, `revenue-agent`, `support-agent` |
| operations_job_status_v2.v1 | calibration | 5 | `inventory-control-agent-holdout`, `job-insights-skill-holdout`, `operations-context-provider-holdout`, `operations-dashboard-agent-holdout`, `support-escalation-agent-holdout` |
| risk_assessment_v2.v1 | primary | 6 | `compliance-archive-control`, `credit-support-agent`, `fraud-signals-control`, `risk-context-provider`, `risk-guidance-skill`, `risk-review-agent` |
| shipment_quote_v2.v1 | primary | 6 | `carrier-scorecard-control`, `fulfillment-insights-skill`, `route-planning-control`, `shipping-context-provider`, `shipping-portal-agent`, `warehouse-assistant-agent` |

## Frozen run order

- `1:billing_summary_v2.v1:r1:mcpatch_discovery`
- `2:billing_summary_v2.v1:r1:code_search_discovery`
- `3:shipment_quote_v2.v1:r2:mcpatch_discovery`
- `4:shipment_quote_v2.v1:r2:code_search_discovery`
- `5:shipment_quote_v2.v1:r3:mcpatch_discovery`
- `6:shipment_quote_v2.v1:r3:code_search_discovery`
- `7:risk_assessment_v2.v1:r3:code_search_discovery`
- `8:risk_assessment_v2.v1:r3:mcpatch_discovery`
- `11:shipment_quote_v2.v1:r1:mcpatch_discovery`
- `12:shipment_quote_v2.v1:r1:code_search_discovery`
- `15:billing_summary_v2.v1:r2:mcpatch_discovery`
- `16:billing_summary_v2.v1:r2:code_search_discovery`
- `17:billing_summary_v2.v1:r3:code_search_discovery`
- `18:billing_summary_v2.v1:r3:mcpatch_discovery`
- `19:risk_assessment_v2.v1:r2:code_search_discovery`
- `20:risk_assessment_v2.v1:r2:mcpatch_discovery`
- `21:risk_assessment_v2.v1:r1:code_search_discovery`
- `22:risk_assessment_v2.v1:r1:mcpatch_discovery`
- `23:billing_summary_v2.v1:r1:repo_only_repair`
- `24:billing_summary_v2.v1:r1:mcpatch_guided_repair`
- `25:billing_summary_v2.v1:r3:repo_only_repair`
- `26:billing_summary_v2.v1:r3:mcpatch_guided_repair`
- `27:shipment_quote_v2.v1:r2:mcpatch_guided_repair`
- `28:shipment_quote_v2.v1:r2:repo_only_repair`
- `29:billing_summary_v2.v1:r2:mcpatch_guided_repair`
- `30:billing_summary_v2.v1:r2:repo_only_repair`
- `31:shipment_quote_v2.v1:r1:mcpatch_guided_repair`
- `32:shipment_quote_v2.v1:r1:repo_only_repair`
- `35:risk_assessment_v2.v1:r1:repo_only_repair`
- `36:risk_assessment_v2.v1:r1:mcpatch_guided_repair`
- `37:shipment_quote_v2.v1:r3:mcpatch_guided_repair`
- `38:shipment_quote_v2.v1:r3:repo_only_repair`
- `39:risk_assessment_v2.v1:r2:repo_only_repair`
- `40:risk_assessment_v2.v1:r2:mcpatch_guided_repair`
- `43:risk_assessment_v2.v1:r3:mcpatch_guided_repair`
- `44:risk_assessment_v2.v1:r3:repo_only_repair`
- `45:shipment_quote_v2.v1:r1:agent_only_end_to_end`
- `46:shipment_quote_v2.v1:r1:mcpatch_end_to_end`
- `49:billing_summary_v2.v1:r3:agent_only_end_to_end`
- `50:billing_summary_v2.v1:r3:mcpatch_end_to_end`
- `51:risk_assessment_v2.v1:r2:agent_only_end_to_end`
- `52:risk_assessment_v2.v1:r2:mcpatch_end_to_end`
- `53:billing_summary_v2.v1:r1:mcpatch_end_to_end`
- `54:billing_summary_v2.v1:r1:agent_only_end_to_end`
- `55:risk_assessment_v2.v1:r1:agent_only_end_to_end`
- `56:risk_assessment_v2.v1:r1:mcpatch_end_to_end`
- `57:shipment_quote_v2.v1:r2:mcpatch_end_to_end`
- `58:shipment_quote_v2.v1:r2:agent_only_end_to_end`
- `59:shipment_quote_v2.v1:r3:agent_only_end_to_end`
- `60:shipment_quote_v2.v1:r3:mcpatch_end_to_end`
- `61:billing_summary_v2.v1:r2:agent_only_end_to_end`
- `62:billing_summary_v2.v1:r2:mcpatch_end_to_end`
- `65:risk_assessment_v2.v1:r3:mcpatch_end_to_end`
- `66:risk_assessment_v2.v1:r3:agent_only_end_to_end`

## Primary condition results

| Condition | n | Successful | First attempt | Within budget | Semantic safe | Governance safe |
|---|---:|---:|---:|---:|---:|---:|
| code_search_discovery | 9 | 0.0% (0/9) | n/a | n/a | n/a | n/a |
| mcpatch_discovery | 9 | 100.0% (9/9) | n/a | n/a | n/a | n/a |
| repo_only_repair | 9 | 100.0% (9/9) | 100.0% (9/9) | 100.0% (9/9) | 100.0% (9/9) | 100.0% (9/9) |
| mcpatch_guided_repair | 9 | 88.9% (8/9) | 66.7% (6/9) | 88.9% (8/9) | 100.0% (8/8) | 100.0% (8/8) |
| agent_only_end_to_end | 9 | 0.0% (0/9) | 0.0% (0/9) | n/a | not observable (0/0) | not observable (0/0) |
| mcpatch_end_to_end | 9 | 100.0% (9/9) | 66.7% (6/9) | n/a | 100.0% (9/9) | 100.0% (9/9) |

## Discovery metrics

| Condition | Set | n | TP | FP | FN | TN | Precision | Recall | F1 | Accuracy | Exact |
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| code_search_discovery | affected consumers | 9 | 8 | 6 | 10 | 0 | 57.1% | 44.4% | 50.0% | 33.3% | 0.0% (0/9) |
| code_search_discovery | patch targets | 9 | 12 | 6 | 6 | 30 | 66.7% | 66.7% | 66.7% | 77.8% | 0.0% (0/9) |
| code_search_discovery | no-change repositories | 9 | 18 | 0 | 18 | 18 | 100.0% | 50.0% | 66.7% | 66.7% | 0.0% (0/9) |
| mcpatch_discovery | affected consumers | 9 | 18 | 0 | 0 | 0 | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% (9/9) |
| mcpatch_discovery | patch targets | 9 | 18 | 0 | 0 | 36 | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% (9/9) |
| mcpatch_discovery | no-change repositories | 9 | 36 | 0 | 0 | 18 | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% (9/9) |

## Repair metrics

| Condition | n | First attempt | Within budget | Attempts total | Mean | Median | Integration | Controls | Changed files | Changed lines |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| repo_only_repair | 9 | 100.0% (9/9) | 100.0% (9/9) | 18 | 2.000 | 2.000 | 100.0% (9/9) | 100.0% (9/9) | 36 | 139 |
| mcpatch_guided_repair | 9 | 66.7% (6/9) | 88.9% (8/9) | 22 | 2.444 | 2.000 | 100.0% (9/9) | 100.0% (9/9) | 44 | 296 |

## End-to-end metrics

| Condition | n | Ready for review | Fleet repaired | Consumers protected | Targets repaired | Correct target + patch | No unnecessary | Semantic safe | Governance safe | Integration | Controls |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| agent_only_end_to_end | 9 | 0.0% (0/9) | 66.7% (6/9) | 22.2% (2/9) | 66.7% (6/9) | 0.0% (0/9) | 33.3% (3/9) | not observable (0/0) | not observable (0/0) | 33.3% (2/6) | 0.0% (0/9) |
| mcpatch_end_to_end | 9 | 100.0% (9/9) | 100.0% (9/9) | 100.0% (9/9) | 100.0% (9/9) | 100.0% (9/9) | 100.0% (9/9) | 100.0% (9/9) | 100.0% (9/9) | 100.0% (9/9) | 100.0% (9/9) |

Raw end-to-end totals:

- `agent_only_end_to_end` raw end-to-end totals: attempts 26; unnecessary repository modifications 6; missing repair targets 6.
- `mcpatch_end_to_end` raw end-to-end totals: attempts 21; unnecessary repository modifications 0; missing repair targets 0.

# Sanitized benchmark prompt-template comparison

The repository-only template supplies contracts, path policy, and approved commands.
The guided template adds verified migration assertions with evidence IDs and checksums.
No rendered prompt, provider transcript, hidden assertion, or answer key is included.

- Repository-only template: `f2396330ea520ec6c6d12f08658bcff82bdd114f7fecd2549b346ec666c199cd`
- Guided template: `baba1e75599f2cb388e6c93fdff9a70747798c3d5bebb0653670d88aea0c4e66`

## Per-scenario results

| Scenario | Tier | Condition | Success | FP repos | FN repos | Unnecessary | Missing | Attempts |
|---|---|---|---:|---:|---:|---:|---:|---:|
| billing_summary_v2.v1 | primary | agent_only_end_to_end | 0.0% (0/3) | 3 | 0 | 3 | 0 | 12 |
| billing_summary_v2.v1 | primary | code_search_discovery | 0.0% (0/3) | 3 | 0 | 0 | 0 | 0 |
| billing_summary_v2.v1 | primary | mcpatch_discovery | 100.0% (3/3) | 0 | 0 | 0 | 0 | 0 |
| billing_summary_v2.v1 | primary | mcpatch_end_to_end | 100.0% (3/3) | 0 | 0 | 0 | 0 | 6 |
| billing_summary_v2.v1 | primary | mcpatch_guided_repair | 100.0% (3/3) | 0 | 0 | 0 | 0 | 8 |
| billing_summary_v2.v1 | primary | repo_only_repair | 100.0% (3/3) | 0 | 0 | 0 | 0 | 6 |
| risk_assessment_v2.v1 | primary | agent_only_end_to_end | 0.0% (0/3) | 2 | 2 | 2 | 2 | 9 |
| risk_assessment_v2.v1 | primary | code_search_discovery | 0.0% (0/3) | 2 | 2 | 0 | 0 | 0 |
| risk_assessment_v2.v1 | primary | mcpatch_discovery | 100.0% (3/3) | 0 | 0 | 0 | 0 | 0 |
| risk_assessment_v2.v1 | primary | mcpatch_end_to_end | 100.0% (3/3) | 0 | 0 | 0 | 0 | 8 |
| risk_assessment_v2.v1 | primary | mcpatch_guided_repair | 66.7% (2/3) | 0 | 0 | 0 | 0 | 6 |
| risk_assessment_v2.v1 | primary | repo_only_repair | 100.0% (3/3) | 0 | 0 | 0 | 0 | 6 |
| shipment_quote_v2.v1 | primary | agent_only_end_to_end | 0.0% (0/3) | 1 | 4 | 1 | 4 | 5 |
| shipment_quote_v2.v1 | primary | code_search_discovery | 0.0% (0/3) | 1 | 4 | 0 | 0 | 0 |
| shipment_quote_v2.v1 | primary | mcpatch_discovery | 100.0% (3/3) | 0 | 0 | 0 | 0 | 0 |
| shipment_quote_v2.v1 | primary | mcpatch_end_to_end | 100.0% (3/3) | 0 | 0 | 0 | 0 | 7 |
| shipment_quote_v2.v1 | primary | mcpatch_guided_repair | 100.0% (3/3) | 0 | 0 | 0 | 0 | 8 |
| shipment_quote_v2.v1 | primary | repo_only_repair | 100.0% (3/3) | 0 | 0 | 0 | 0 | 6 |

### Per-scenario discovery and end-to-end detail

- `billing_summary_v2.v1` / `agent_only_end_to_end`: consumer P/R/F1 57.1%/66.7%/61.5% (TP/FP/FN/TN 4/3/2/0); target P/R/F1 66.7%/100.0%/80.0% (TP/FP/FN/TN 6/3/0/9); control accuracy 83.3%; ready 0.0% (0/3); correct target + patch 0.0% (0/3).
- `billing_summary_v2.v1` / `code_search_discovery`: consumer P/R/F1 57.1%/66.7%/61.5% (TP/FP/FN/TN 4/3/2/0); target P/R/F1 66.7%/100.0%/80.0% (TP/FP/FN/TN 6/3/0/9); control accuracy 83.3%.
- `billing_summary_v2.v1` / `mcpatch_discovery`: consumer P/R/F1 100.0%/100.0%/100.0% (TP/FP/FN/TN 6/0/0/0); target P/R/F1 100.0%/100.0%/100.0% (TP/FP/FN/TN 6/0/0/12); control accuracy 100.0%.
- `billing_summary_v2.v1` / `mcpatch_end_to_end`: consumer P/R/F1 100.0%/100.0%/100.0% (TP/FP/FN/TN 6/0/0/0); target P/R/F1 100.0%/100.0%/100.0% (TP/FP/FN/TN 6/0/0/12); control accuracy 100.0%; ready 100.0% (3/3); correct target + patch 100.0% (3/3).
- `risk_assessment_v2.v1` / `agent_only_end_to_end`: consumer P/R/F1 60.0%/50.0%/54.5% (TP/FP/FN/TN 3/2/3/0); target P/R/F1 66.7%/66.7%/66.7% (TP/FP/FN/TN 4/2/2/10); control accuracy 66.7%; ready 0.0% (0/3); correct target + patch 0.0% (0/3).
- `risk_assessment_v2.v1` / `code_search_discovery`: consumer P/R/F1 60.0%/50.0%/54.5% (TP/FP/FN/TN 3/2/3/0); target P/R/F1 66.7%/66.7%/66.7% (TP/FP/FN/TN 4/2/2/10); control accuracy 66.7%.
- `risk_assessment_v2.v1` / `mcpatch_discovery`: consumer P/R/F1 100.0%/100.0%/100.0% (TP/FP/FN/TN 6/0/0/0); target P/R/F1 100.0%/100.0%/100.0% (TP/FP/FN/TN 6/0/0/12); control accuracy 100.0%.
- `risk_assessment_v2.v1` / `mcpatch_end_to_end`: consumer P/R/F1 100.0%/100.0%/100.0% (TP/FP/FN/TN 6/0/0/0); target P/R/F1 100.0%/100.0%/100.0% (TP/FP/FN/TN 6/0/0/12); control accuracy 100.0%; ready 100.0% (3/3); correct target + patch 100.0% (3/3).
- `shipment_quote_v2.v1` / `agent_only_end_to_end`: consumer P/R/F1 50.0%/16.7%/25.0% (TP/FP/FN/TN 1/1/5/0); target P/R/F1 66.7%/33.3%/44.4% (TP/FP/FN/TN 2/1/4/11); control accuracy 50.0%; ready 0.0% (0/3); correct target + patch 0.0% (0/3).
- `shipment_quote_v2.v1` / `code_search_discovery`: consumer P/R/F1 50.0%/16.7%/25.0% (TP/FP/FN/TN 1/1/5/0); target P/R/F1 66.7%/33.3%/44.4% (TP/FP/FN/TN 2/1/4/11); control accuracy 50.0%.
- `shipment_quote_v2.v1` / `mcpatch_discovery`: consumer P/R/F1 100.0%/100.0%/100.0% (TP/FP/FN/TN 6/0/0/0); target P/R/F1 100.0%/100.0%/100.0% (TP/FP/FN/TN 6/0/0/12); control accuracy 100.0%.
- `shipment_quote_v2.v1` / `mcpatch_end_to_end`: consumer P/R/F1 100.0%/100.0%/100.0% (TP/FP/FN/TN 6/0/0/0); target P/R/F1 100.0%/100.0%/100.0% (TP/FP/FN/TN 6/0/0/12); control accuracy 100.0%; ready 100.0% (3/3); correct target + patch 100.0% (3/3).

## Failure modes

| Condition | Layer | Code | Count |
|---|---|---|---:|
| agent_only_end_to_end | composition | affected_consumer_mismatch | 9 |
| agent_only_end_to_end | composition | missing_repair_target | 6 |
| agent_only_end_to_end | composition | no_change_repository_mismatch | 9 |
| agent_only_end_to_end | composition | unnecessary_repair_target | 6 |
| agent_only_end_to_end | discovery | malformed_discovery_json | 3 |
| agent_only_end_to_end | integration | selected_no_change_repository_modified | 12 |
| agent_only_end_to_end | patch_policy | policy_diff_too_large | 2 |
| code_search_discovery | discovery | malformed_discovery_json | 3 |
| mcpatch_end_to_end | contract | contract_required_output | 3 |
| mcpatch_guided_repair | contract | contract_required_output | 2 |
| mcpatch_guided_repair | contract | evaluation_plugin_unavailable | 2 |
| mcpatch_guided_repair | patch_policy | policy_diff_too_large | 1 |

## Retained infrastructure-invalid attempts

No infrastructure-invalid attempts were retained.

## Timing and model-usage coverage

- `code_search_discovery` elapsed seconds: coverage 9/9; total 829.788; mean 92.199; median 92.911; min 70.699; max 114.968.
  Input tokens: coverage 9/9; total 1371224.000; mean 152358.222; median 157523.000; min 74944.000; max 188284.000. Output tokens: coverage 9/9; total 28567.000; mean 3174.111; median 3174.000; min 2695.000; max 3546.000. Cost USD: not observable (0/9).
- `mcpatch_discovery` elapsed seconds: coverage 9/9; total 0.000; mean 0.000; median 0.000; min 0.000; max 0.000.
  Input tokens: not observable (0/9). Output tokens: not observable (0/9). Cost USD: not observable (0/9).
- `repo_only_repair` elapsed seconds: coverage 9/9; total 1250.608; mean 138.956; median 136.857; min 102.922; max 189.727.
  Input tokens: coverage 9/9; total 2559434.000; mean 284381.556; median 286978.000; min 238312.000; max 316311.000. Output tokens: coverage 9/9; total 23420.000; mean 2602.222; median 2427.000; min 2277.000; max 3619.000. Cost USD: not observable (0/9).
- `mcpatch_guided_repair` elapsed seconds: coverage 9/9; total 2018.999; mean 224.333; median 204.309; min 151.329; max 371.797.
  Input tokens: coverage 9/9; total 3985045.000; mean 442782.778; median 380150.000; min 273602.000; max 758193.000. Output tokens: coverage 9/9; total 43570.000; mean 4841.111; median 4325.000; min 2980.000; max 8155.000. Cost USD: not observable (0/9).
- `agent_only_end_to_end` elapsed seconds: coverage 9/9; total 2890.552; mean 321.172; median 385.345; min 70.699; max 508.738.
  Input tokens: coverage 9/9; total 5380631.000; mean 597847.889; median 677949.000; min 104097.000; max 1005255.000. Output tokens: coverage 9/9; total 69236.000; mean 7692.889; median 8747.000; min 2695.000; max 12100.000. Cost USD: not observable (0/9).
- `mcpatch_end_to_end` elapsed seconds: coverage 9/9; total 1760.738; mean 195.638; median 167.955; min 143.445; max 287.717.
  Input tokens: not observable (0/9). Output tokens: not observable (0/9). Cost USD: not observable (0/9).

## Baseline matches or wins

- `billing_summary_v2.v1:r1:repair:baseline=1:mcpatch=1:equal`
- `billing_summary_v2.v1:r2:repair:baseline=1:mcpatch=1:equal`
- `billing_summary_v2.v1:r3:repair:baseline=1:mcpatch=1:equal`
- `risk_assessment_v2.v1:r1:repair:baseline=1:mcpatch=0:baseline_better`
- `risk_assessment_v2.v1:r2:repair:baseline=1:mcpatch=1:equal`
- `risk_assessment_v2.v1:r3:repair:baseline=1:mcpatch=1:equal`
- `shipment_quote_v2.v1:r1:repair:baseline=1:mcpatch=1:equal`
- `shipment_quote_v2.v1:r2:repair:baseline=1:mcpatch=1:equal`
- `shipment_quote_v2.v1:r3:repair:baseline=1:mcpatch=1:equal`

## MCPatch failures

- `trial-1771afcd1573c74c32b6`

## Limitations

These small synthetic Python scenarios do not establish production-scale or cross-language generalization. Percentages always retain their raw counts; unavailable timing, token, model, and cost observations are not imputed.

## Conclusion

Within the scored synthetic Python scenarios, MCPatch achieved a higher paired ready-for-review rate; this evidence does not establish performance outside the supported scenarios.
