Science Explained/Brief
Vendor-Native Coding Harnesses Show No Clear Average Advantage in Paired Test
A new arXiv paper compares agentic coding harnesses paired with the same models on a private, contamination-controlled suite. It reports no resolved average advantage for either harness, with opposite results across task types and unresolved cost ordering.
BriefPublished 14 September 20261 min read1 linked source · 8 checked facts
The study ran 80 tasks under claude-agent-sdk and deepagents on claude-opus-4-8, and under openai-codex SDK and deepagents on gpt-5.5. 792 of 800 planned runs were graded by an isolated oracle. For Opus 4.8, the average difference was -1.25 percentage points (48.8% vs 50.0%, 95% CI [-10.0, +7.5]); for GPT-5.5, +1.25 pp (55.6% vs 54.4%, CI [-4.4, +6.9]). The Opus average hid opposite strata: native trailed by 9.0 pp on 61 repository tasks and led by 23.7 pp on 19 contest tasks, a partition chosen after seeing the data. Cost per solved task favored the neutral harness in observed usage, but missing usage records leave the billed ordering unresolved.
Our view
The result undercuts a simple vendor-native advantage story, but the task-type split and cost uncertainty mean it is not a settled verdict.
What the reporting says: Neither same-model harness contrast resolves an average advantage, with CIs crossing zero, and that the Opus average combines opposite strata whose partition was chosen after seeing the data and needs a designed replication.