Agentic Harnesses for Coding Agents
What does the harness — the scaffold, tool surface, control loop, context manager, verifier, sandbox and orchestrator around a language model — actually contribute to a coding agent? A state-of-the-art digest across the academic literature and primary industry sources, as of 10 September 2026.
The one-paragraph answer
The harness has become a first-class, independently studied engineering object. Three source-code taxonomies published in 2026 converge on roughly the same decomposition, industry now puts the word "harness" in the titles of its engineering posts, and a small experimental literature holds the model fixed and varies the harness, or vice versa. That literature disagrees: the largest crossed study finds a 23.8-point spread across harnesses at fixed model and task; a longitudinal study of 35 harness releases at fixed model finds no significant change in resolve rate but a doubling of cost; a 9,374-trajectory study finds the model dominates and the framework gap shrinks each model generation. The vendors that train models on their own harness and the vendors that call a harness "an encoding of what the model can't do yet" are describing the same thing from opposite ends: harness and model are being co-designed, so their contributions are decreasingly separable from outside.
Dataset at a glance
| Records | 268 — 179 academic, 70 industry, 19 open-source |
|---|---|
| By year | 2026: 134 · 2025: 86 · 2024: 39 · 2023: 7 |
| Empirical or benchmark evidence | 176 records |
| Speak to model-vs-harness attribution | 57 records |
| Surfaced by Consensus | 131 records (105 found only there) |
| Independently re-verified | 83 confirmed, 15 minor metadata fixes, 0 wrong, 0 not found |
Method
Four channels were run on the same day: web and arXiv search for the harness-design literature; the same for evaluation harnesses and benchmarks; primary industry sources (Anthropic, OpenAI, Microsoft and GitHub, Google, Cognition, Cursor, Factory, All Hands, the SWE-agent group, Augment, Replit, Amazon, Sourcegraph); and eight Consensus Deep searches whose cited papers were resolved back to primary pages. Records were written only from pages actually fetched. A separate, adversarially prompted verification pass then re-fetched 102 of the records, including every record whose number is quoted in the synthesis, and its corrections were applied to the dataset.