arXiv:2607. 24882v1 Announce Type: cross Abstract: Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed for the task.
By Bowen Qin, Yi Xie
arXiv:2607. 25066v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate reasoning traces, actions, and tool observations that can eventually exceed a model's fixed context window.
By Thang Dang, Yuma Ichikawa, Sakina Fatima, Koichi Shirahata
arXiv:2607. 16848v1 Announce Type: cross Abstract: Long-term memory is becoming a core component of LLM agents, but most memory benchmarks evaluate conversations or compact summaries, while research agents need to restore evidence from full scientific papers.
By Maksim Sheverev, David Finkelstein, Sergey Nikolenko
The paper introduces State‑Conditioned Minimal Sufficient Evidence Recovery (SER), a method that, given a coding agent’s current state, reconstructs a compact set of evidence passages that collectively provide all facts needed for the agent’s next decision. Using the SERBench dataset of 500 held‑out states from 45 repositories, the authors show that their MSS‑Complement approach recovers a complete evidence set for 73.0 % of states with five items and 80.6 % with eight, outperforming baseline ranking methods. The study also demonstrates that this set‑level policy improves downstream performance on AMA‑Bench and highlights the importance of retrieving missing facts rather than merely re‑ranking similar passages.
By Zhexi Feng, Ruiyi Zhang, Yongbo Yang, Pengtao Xie
The paper introduces MERIT, a benchmark that evaluates the marginal benefit of long‑term memory for tool‑using large language model agents while explicitly accounting for cost. MERIT provides episodic tool‑use tasks across three domains, verifies dependence on earlier‑episode facts, and measures memory operations in tokens and dollars. Experiments on GPT‑4.1‑mini, Claude Haiku 4.5, and Claude Sonnet 5 show that memory can significantly improve task success, but its utility varies widely across models and memory implementations, and full replay is rarely cost‑effective.
By Shweta Mishra, Shashank Mishra
arXiv:2609.22223v1 Announce Type: cross
Abstract: Long-form factuality verification is commonly implemented as a static decompose-search-verify pipeline, with separately prompted modules processing c...
By Kening Zheng, Aoying Zheng, Zhigang Chang, Yazhi Guo, Miaotian Guo, Qingwei Zong, Xianhai Xie, Weiqiang Jin, Chengze Li, Hanrong Zhang, Jie Yang, Wei-Chieh Huang, Lingzhe Zhang, Liancheng Fang, Xin Zou, Hanqian Li, Jiahao Huo, Yibo Yan, Zizhuang Deng, Lei Miao, Wei Guo, Haihong Tang, Bo Zheng, Philip S. Yu
arXiv:2606. 29914v1 Announce Type: cross Abstract: Agent memory systems are increasingly evaluated against RAG and full-context baselines, but reported gains often mix changes in the memory method with changes in the language model, embedding model, or retrieval pipeline, making it unclear what is actually being measured.
By Kuan Wang
The paper introduces a Channel‑Boosted Multi‑Agent System (CB‑MAS), instantiated as IC‑MAS, to classify document sensitivity without the input‑length truncation problem of standard transformers. IC‑MAS uses a Channel Critic Agent to adaptively weight two first‑window encoders and Consultation Agents to exchange belief states, achieving 90.72% accuracy and 91.23% F1‑score while reducing computation by ~54% compared to a fixed‑round baseline. The authors provide explainability via LIME/SHAP, multi‑agent evaluation, and a transparent discussion of limitations.
By Aleesha Zainab, Asifullah Khan, Muhammad Ahmed Khalid, Faheem Ullah Khan
arXiv:2607. 28609v2 Announce Type: replace Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world.
By Qiushi Sun, Kanzhi Cheng, Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu, Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong
The paper reports the ABAI submission to COLIEE 2026 Task 1, a case law retrieval challenge that suppresses cited passages, and details a four‑stage retrieval pipeline: multi‑view BM25 with reciprocal rank fusion, neural reranking, graph‑based features via a graph attention network, and a LightGBM meta‑learner over 34 features. The best run achieved an F1 score of 0.177 on the official test set, compared to a cross‑validated 0.311, and the authors attribute the gap to a recall ceiling, temporal distribution shift, and threshold miscalibration. A controlled post‑hoc study examined the impact of threshold transfer, decision quality across time, and query similarity, and identified specific remedies—such as BM25 length‑normalisation tuning, event‑triple views, and dense fusion—that improved recall, while other interventions had no effect.
By Minhan Cho, Soyoung Park, Daejin Choi, Jinyoung Han
arXiv:2609.38021v1 Announce Type: cross
Abstract: We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, cove...
By Christopher J. Chanhnourack
arXiv:2609.26086v1 Announce Type: new
Abstract: An agentic retrieval system issues a sequence of search queries and must decide, at each step, whether the evidence collected so far is enough to stop....
By Daeyoung Roh, Donghee Han