The research turn toward forgetting — and the survey that shows how much is still proposal
This month's agent-memory papers are unusually well-posed. A survey published alongside them makes the field's real problem visible: the ratio of architecture proposals to replicated results is not healthy.
A survey of multimodal agent AI catalogues the subfield's advances, and the most useful thing it does is make a ratio legible: a great deal of proposed architecture, comparatively little replicated finding. That is characteristic of a field moving faster than its evaluation methodology.
Where the testable claims are
MemCtrl's learned retention gate and AMA's consistency verification are refreshing precisely because they can be wrong in specific ways. A gate either improves downstream task performance under a fixed budget or it does not. Verification either catches contradictory memories or it does not. Falsifiability is the scarce resource here.
The benchmark that resists the method
Coding agents turning toward ARC-AGI-3 is the other healthy sign. ARC's value has always been adversarial: each version is redesigned after the previous one is solved in ways its authors found unconvincing. Progress on a benchmark built to resist your dominant method is informative in a way progress on a saturated one is not.
Applying coding agents rather than raw models is the notable shift — it reframes the question from whether the model sees the pattern to whether a system can search for a program producing it. That is arguably what the benchmark was always asking.
The standard worth holding
A field that publishes frameworks faster than it validates them accumulates vocabulary rather than knowledge. The memory strand is currently the exception. It deserves the replication effort more than the next architecture does.
JCST — Multimodal Agent AI: A Survey of Recent Advances and Future Directions → · VoltAgent — Awesome AI Agent Papers — 2026 collection →