Memory for Coding Agents
What persists when a coding agent's session ends — repository facts, preferences, episodes, skills, rules — and who gets to write it? A state-of-the-art digest across the academic literature and primary industry sources, with particular attention to human-in-the-loop governance, as of 14 September 2026.
The one-paragraph answer
The field has converged on a vocabulary — memory types × memory operations, with memory, skills and rules understood as one artifact at different compression ratios — but not on evidence. The best-controlled studies find that persistent memory does not raise coding-agent success rates by default: repository context files do not generally improve task success while adding over 20% inference cost, random rules match expert-curated ones, 39 of 49 real software-engineering skills give zero gain, and a pre-registered five-way memory comparison finds gains only when the new problem is a near-copy of a stored one. What does move the number is validation of the write — a tested-before-accept gate turns "no better than no skills" into a 48-point gain — and delivery of the read: agents left to choose make zero voluntary memory operations in 114 turns. Industry has reached the same conclusion from the other side: the one vendor with an A/B on a real engineering outcome gates memories with citation checks and an expiry; two vendors retrofitted per-memory human approval after shipping silent auto-memory; one retired auto-memory for human-authored skills. Nobody has yet run the experiment that matters for a governance stance: human-ratified versus execution-validated memory promotion, head to head, on the same repository tasks.
Dataset at a glance
| Records | 283 — 197 academic, 62 industry, 24 open-source |
|---|---|
| By year | 2026: 215 · 2025: 50 · 2024: 10 · 2023: 7 |
| Human in the loop somewhere | 131 records; 20 with pre-write approval; 44 human-authored memory |
| Write validated at all | 132 records |
| Benchmarks / security & privacy | 56 / 52 records |
| Surfaced by Consensus | 220 records |
| Independently re-verified | 100 confirmed, 8 minor metadata fixes, 4 with a number or page not found, 0 wrong |
Method
Four channels on the same day: a literature pass on memory mechanisms and taxonomies; a pass on benchmarks, failure modes and security; primary industry sources for every major coding agent and memory platform; and a human-in-the-loop and governance pass. The Consensus connector was the primary academic discovery channel, with arXiv abstract pages fetched for every number. A separate adversarially prompted verification agent re-fetched 112 of the records — every low- or medium-confidence record, every number quoted in the research memos, plus a random thirty — and its corrections were applied. Every record is coded on the paper's memory scheme (scope, storage, write policy, validation, evidence) plus a human-in-the-loop field, so the explorer can be filtered on exactly the questions the synthesis asks.