← SoA Agentic SE

Agentic Harnesses for Coding Agents

What does the harness — the scaffold, tool surface, control loop, context manager, verifier, sandbox and orchestrator around a language model — actually contribute to a coding agent? A state-of-the-art digest across the academic literature and primary industry sources, as of 10 September 2026.

Read the synthesisEight sections: taxonomy, the design space dimension by dimension, industry landscape, evaluation harnesses and their validity, field evidence, the attribution question, gaps. Explore the 268 sourcesFilter by sector, year, category, evidence type, discovery channel and verification status. Every claim in the synthesis links here. Consensus deep searchesThe eight consensus.app "Deep" queries, their generated syntheses, and the papers each surfaced.

The one-paragraph answer

The harness has become a first-class, independently studied engineering object. Three source-code taxonomies published in 2026 converge on roughly the same decomposition, industry now puts the word "harness" in the titles of its engineering posts, and a small experimental literature holds the model fixed and varies the harness, or vice versa. That literature disagrees: the largest crossed study finds a 23.8-point spread across harnesses at fixed model and task; a longitudinal study of 35 harness releases at fixed model finds no significant change in resolve rate but a doubling of cost; a 9,374-trajectory study finds the model dominates and the framework gap shrinks each model generation. The vendors that train models on their own harness and the vendors that call a harness "an encoding of what the model can't do yet" are describing the same thing from opposite ends: harness and model are being co-designed, so their contributions are decreasingly separable from outside.

Dataset at a glance

Records268 — 179 academic, 70 industry, 19 open-source
By year2026: 134 · 2025: 86 · 2024: 39 · 2023: 7
Empirical or benchmark evidence176 records
Speak to model-vs-harness attribution57 records
Surfaced by Consensus131 records (105 found only there)
Independently re-verified83 confirmed, 15 minor metadata fixes, 0 wrong, 0 not found

Method

Four channels were run on the same day: web and arXiv search for the harness-design literature; the same for evaluation harnesses and benchmarks; primary industry sources (Anthropic, OpenAI, Microsoft and GitHub, Google, Cognition, Cursor, Factory, All Hands, the SWE-agent group, Augment, Replit, Amazon, Sourcegraph); and eight Consensus Deep searches whose cited papers were resolved back to primary pages. Records were written only from pages actually fetched. A separate, adversarially prompted verification pass then re-fetched 102 of the records, including every record whose number is quoted in the synthesis, and its corrections were applied to the dataset.