Useful Automation/Brief
Harness choice shows no resolved average task advantage in paired agentic coding tests
A preprint reports paired same-model contrasts on a private contamination-controlled suite of 256 repository and post-cutoff contest tasks. It finds no resolved average harness advantage, with cost and completion caveats.
BriefPublished 14 September 20261 min read1 linked source · 5 checked facts
The study ran the same 80 tasks under claude-agent-sdk and deepagents on claude-opus-4-8, and under openai-codex SDK and deepagents on gpt-5.5. Neither contrast resolved an average advantage: -1.25 pp for Opus 4.8 and +1.25 pp for GPT-5.5. The Opus average combines opposite strata: the native harness trails by 9.0 pp on the 61 repository tasks and leads by 23.7 pp on the 19 contest tasks, a partition chosen after seeing the data and needing a designed replication. Also, 22 of 81 runs cancelled at the wall-clock ceiling had produced a passing patch. Re-priced from raw per-turn usage at frozen list prices, the neutral harness cost 1.3 to 1.6 times as much per solved task on Opus 4.8 and 1.2 times on GPT-5.5, though billed ordering is unresolved because 58 Anthropic runs left no usage record.
Our view
Teams should not assume vendor-native harnesses improve task success and should measure cost per solved task and completion behavior per task type.
What the reporting says: Neither contrast resolves an average advantage for either harness, that the Opus average combines opposite repository and contest strata chosen after seeing the data, and that 22 of 81 wall-clock-cancelled runs had produced a passing patch.