arXiv:2606. 14470v1 Announce Type: new Abstract: Large language model (LLM) reasoning is ephemeral: chains of thought vanish with the context window, pruned search branches leave no record, and memory buffers cannot be diffed, merged, or audited.
By Pavan C Shekar, Abhishek H S, Aswanth Krishnan
arXiv:2609.23570v1 Announce Type: cross
Abstract: Coding agents operate on real repository coding tasks, and persistent memory systems promise to reuse experience across tasks. Yet existing evaluatio...
By Liyang Fan, Yingcheng Shi, Yongbin Li, Chenghao Sun, Xin Chen, Xander Xu, Hu Wei, Shiwen Ni, Min Yang, Jieping Ye
arXiv:2606. 09900v1 Announce Type: cross Abstract: Long-term memory is the missing layer for LLM agents: across sessions they forget, and the common workaround -- replaying the whole history into the prompt -- is expensive, slow, and, as distractors accumulate, less accurate.
By Liuyin Wang
arXiv:2606. 11976v1 Announce Type: cross Abstract: Software engineering tools increasingly rely on LLM based agents to localize files to change to resolve a software issue.
By Akeela Darryl Fattha, Kia Ying Chua, Lingxiao Jiang, Laura Wynter
Agora is a system that uses Git as a shared memory for autonomous research agents, recording each claim as an immutable commit in an append‑only directed acyclic graph. In a 12‑day run, 13 language‑model workers independently explored a weight‑transfer problem, producing 1,703 contributions that improved a 119.6M‑parameter model’s performance from 3.39 to 1.899 bits per byte. The system’s design includes a diversity‑aware selection rule and an index that tracks the frontier, neglected branches, and verification status of each claim.
By Yifan Zhang, Yunheng Zou, Shaokun Zhang, Jian Hu, Hao Zhang, Binfeng Xu, Jan Kautz, Yi Dong
The paper introduces MERIT, a benchmark that evaluates the marginal benefit of long‑term memory for tool‑using large language model agents while explicitly accounting for cost. MERIT provides episodic tool‑use tasks across three domains, verifies dependence on earlier‑episode facts, and measures memory operations in tokens and dollars. Experiments on GPT‑4.1‑mini, Claude Haiku 4.5, and Claude Sonnet 5 show that memory can significantly improve task success, but its utility varies widely across models and memory implementations, and full replay is rarely cost‑effective.
By Shweta Mishra, Shashank Mishra
Mnemon is a memory agent that stores conversations as raw, dated records and uses a fast System 1 decision model (Jev) to quickly judge the relevance of records, while a slow System 2 LLM plans searches and composes answers. The agent consolidates records into topic timelines and value histories in the background, enabling efficient retrieval without rewriting conversations into structured formats. Experiments show Mnemon achieving high scores on LoCoMo and LongMemEval‑S with low context length and cost, and Jev outperforming LLMs in evidence separation and speed.
By Guangren Wang
arXiv:2608. 05886v1 Announce Type: cross Abstract: Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it.
By Wuya Chen, Yihao yang, Yang Cao, Yue Lin
The paper introduces “Revise”, a runtime system that performs validity-guided, fine-grained recovery for online revisions in structured agent workflows. When a revision arrives, Revise intersects the change with recorded data and control dependencies, propagates the impact through the partially executed DAG, stops invalid work, preserves unaffected progress, and recomputes only the affected region. Experiments on real coding‑agent traces and LangGraph/LLMCompiler applications show that Revise matches a latest‑version oracle, reduces model calls by up to 56%, and improves service‑level objective goodput under load.
By Ruoling Qi, Xuaner Wu, Penghang Liu, Jian Chen, Yirui Liu
The paper introduces environment‑probing curation, a deployment‑compatible method that equips asynchronous curator agents with read‑only world tools to verify, scope, and refresh candidate memories without retraining models. In a GitHub Copilot‑based harness, this approach improves pass rates on CLBench from 39% to 73%, boosts reward metrics, and reduces both query counts and task‑agent costs. Across six APEX management‑consulting tasks, the method consistently outperforms baselines, yielding higher rewards and fewer tool calls while maintaining a compact task‑time interface.
By Susheel Suresh, Hazel Mak, Sahil Bhatnagar, Chhaya Methani, Alejandro Gutierrez Munoz
Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent averages 23 rounds and 631K tokens per resolved issue, with many calls spent on grep, glob, and view_file during repository exploration.
arXiv:2607. 24882v1 Announce Type: cross Abstract: Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed for the task.
By Bowen Qin, Yi Xie