← Research

State of the Art in Agentic Software Engineering

A paper project asking what the field of agentic software engineering has actually built, what has been measured about it, and what would have to be true for published results to transfer. The central claim under test: the field has moved to vertically integrated model–harness products, so attributing outcomes to the model versus the scaffold around it is structurally hard to do from outside.

The research tracks below are reference stores gathered alongside the paper's frozen search protocol, not entries in its evidence base. Each has a synthesis whose every claim links to a source record, a filterable source explorer, and the Consensus search transcripts behind it. Records were written only from pages actually fetched and were re-checked by a separate adversarially prompted verification pass; the explorer shows the verdict per record.

Tracks

Agentic harnesses What the scaffold around a language model contributes: taxonomy, design space, industry landscape, evaluation validity, and the model-vs-harness attribution evidence. 268 sources · 10 September 2026 Agent memory Persistent memory for coding agents: who writes it, when, how it is validated and forgotten, what has been measured, and how humans stay in the loop. 283 sources · 14 September 2026

How the tracks relate

Memory is one of the harness's design dimensions — the context and memory layer, plus the externally authored configuration surface (CLAUDE.md, AGENTS.md, rules, skills) that programs the runtime's behaviour. The harness track covers in-context compaction; the memory track covers what persists across sessions and people. Citations in the memory synthesis that point into the harness store are prefixed harness:.