Phase 5.1 Benchmark Validated

Committed result bytes passed checksum and typed-artifact verification before display.

5764275dba0fe953

Phase 5.1 · Comparative benchmark

Does organizational context help the same coding agent repair the right code safely?

The benchmark pairs one real Codex baseline with MCPatch and DataHub across discovery, repository repair, and the composed end-to-end workflow. Percentages retain raw counts, and unavailable observations remain unavailable.

Methodology

Each pair uses the same provider, model configuration, repository snapshot, contracts, evaluator, attempt budget, timeout, and resource limits. Ground truth and evaluator code remain outside the coding provider.

TrackResearch questionBaselineMCPatch condition
DiscoveryWhich agents are affected, and which repositories actually require repair?Code-search discoveryDeterministic MCPatch and DataHub discovery
RepairDoes DataHub-derived context help the same agent produce a safe patch?Repository-only repairMCPatch-guided repair
End to endDoes the complete workflow reach a ready-for-review fleet repair?Agent-only discovery and repairMCPatch discovery and guided repair

Sanitized context difference

This comparison describes the fixed informational treatment without publishing rendered provider material, hidden requirements, or transcripts.

Repository-only condition

Baseline receives

Contracts, the target repository, allowed and protected paths, and approved repository-native commands. No organizational context.

Template f2396330ea520ec6c6d12f08658bcff82bdd114f7fecd2549b346ec666c199cd

MCPatch-guided condition

Treatment adds

The same inputs plus deterministic migration assertions derived from verified dependency, business-semantic, classification, governance, and ownership evidence. Every assertion carries an evidence ID and checksum.

Template baba1e75599f2cb388e6c93fdff9a70747798c3d5bebb0653670d88aea0c4e66

Scenario tiers

Only primary scenarios may contribute to headline generalization results. Calibration scenarios were already public during development and are reported separately.

Headline cohort

Primary no-gold scenarios

  • billing_summary_v2.v1Primary
  • risk_assessment_v2.v1Primary
  • shipment_quote_v2.v1Primary

Excluded from headline score

Calibration scenarios

Summary metrics

Primary condition rates are generated from immutable raw rows. Semantic and governance measures apply only to repair-bearing tracks.

54 primary trials · 0 calibration trials

ConditionnSuccessfulFirst attemptSemantic safeGovernance safe
Code-search discovery90.0% (0/9)Not applicableNot applicableNot applicable
MCPatch discovery9100.0% (9/9)Not applicableNot applicableNot applicable
Repository-only repair9100.0% (9/9)100.0% (9/9)100.0% (9/9)100.0% (9/9)
MCPatch-guided repair988.9% (8/9)66.7% (6/9)100.0% (8/8)100.0% (8/8)
Agent-only end to end90.0% (0/9)0.0% (0/9)Not observable (0/0)Not observable (0/0)
MCPatch end to end9100.0% (9/9)66.7% (6/9)100.0% (9/9)100.0% (9/9)

Per-scenario comparison

False-positive, false-negative, and unnecessary-repository counts stay visible beside successful outcomes.

ScenarioTierConditionSuccessFP reposFN reposUnnecessary
billing_summary_v2.v1primaryAgent-only end to end0.0% (0/3)303
billing_summary_v2.v1primaryCode-search discovery0.0% (0/3)300
billing_summary_v2.v1primaryMCPatch discovery100.0% (3/3)000
billing_summary_v2.v1primaryMCPatch end to end100.0% (3/3)000
billing_summary_v2.v1primaryMCPatch-guided repair100.0% (3/3)000
billing_summary_v2.v1primaryRepository-only repair100.0% (3/3)000
risk_assessment_v2.v1primaryAgent-only end to end0.0% (0/3)222
risk_assessment_v2.v1primaryCode-search discovery0.0% (0/3)220
risk_assessment_v2.v1primaryMCPatch discovery100.0% (3/3)000
risk_assessment_v2.v1primaryMCPatch end to end100.0% (3/3)000
risk_assessment_v2.v1primaryMCPatch-guided repair66.7% (2/3)000
risk_assessment_v2.v1primaryRepository-only repair100.0% (3/3)000
shipment_quote_v2.v1primaryAgent-only end to end0.0% (0/3)141
shipment_quote_v2.v1primaryCode-search discovery0.0% (0/3)140
shipment_quote_v2.v1primaryMCPatch discovery100.0% (3/3)000
shipment_quote_v2.v1primaryMCPatch end to end100.0% (3/3)000
shipment_quote_v2.v1primaryMCPatch-guided repair100.0% (3/3)000
shipment_quote_v2.v1primaryRepository-only repair100.0% (3/3)000

Failure modes

Failures are part of the result, grouped by condition, evaluator layer, and stable code.

ConditionLayerCodeCount
Agent-only end to endcompositionaffected_consumer_mismatch9
Agent-only end to endcompositionmissing_repair_target6
Agent-only end to endcompositionno_change_repository_mismatch9
Agent-only end to endcompositionunnecessary_repair_target6
Agent-only end to enddiscoverymalformed_discovery_json3
Agent-only end to endintegrationselected_no_change_repository_modified12
Agent-only end to endpatch_policypolicy_diff_too_large2
Code-search discoverydiscoverymalformed_discovery_json3
MCPatch end to endcontractcontract_required_output3
MCPatch-guided repaircontractcontract_required_output2
MCPatch-guided repaircontractevaluation_plugin_unavailable2
MCPatch-guided repairpatch_policypolicy_diff_too_large1

Verified charts

Accessible SVG files are deterministically regenerated from the same typed result document.

Verified benchmark chart: Discovery precision and recall.
Discovery precision and recall
Verified benchmark chart: Repair acceptance rate.
Repair acceptance rate
Verified benchmark chart: Failure modes.
Failure modes
Verified benchmark chart: Unnecessary repositories.
Unnecessary repositories
Verified benchmark chart: Attempts to ready.
Attempts to ready

Evidence and downloads

Only the explicit public result allowlist is downloadable. Runtime homes, credentials, hidden truth, evaluator code, provider transcripts, and private mappings are excluded.

Benchmark version
phase5-1.v1
Specification checksum
f716e1bb453857b7bc34f55f05dac6e0a827c163ecf51f015c03c8896081d93b
SHA256SUMS checksum
5764275dba0fe9531af356efa3e71df334f1a3a6d9af6565169c9f7be4d494a2
Result state
Checksum verified and typed

Limitations

Scope boundaries are part of the evidence, not fine print.

Generated conclusion. Within the scored synthetic Python scenarios, MCPatch achieved a higher paired ready-for-review rate; this evidence does not establish performance outside the supported scenarios.

Replay is read only

Benchmark replay makes no GitHub write, opens no pull request, merges nothing, and deploys nothing. It reconstructs published metrics from committed evidence without loading credentials or contacting external systems.