arXiv AI By Xinyu Guan, Qianyang Zhao, Yuming Deng

Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents

Read the original on arXiv AI →

arXiv:2606. 08151v1 Announce Type: new Abstract: Tool-using LLM agents often fail not because relevant text is absent, but because decisive evidence is not selected, compressed, or surfaced at action time.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

The paper introduces State‑Conditioned Minimal Sufficient Evidence Recovery (SER), a method that, given a coding agent’s current state, reconstructs a compact set of evidence passages that collectively provide all facts needed for the agent’s next decision. Using the SERBench dataset of 500 held‑out states from 45 repositories, the authors show that their MSS‑Complement approach recovers a complete evidence set for 73.0 % of states with five items and 80.6 % with eight, outperforming baseline ranking methods. The study also demonstrates that this set‑level policy improves downstream performance on AMA‑Bench and highlights the importance of retrieving missing facts rather than merely re‑ranking similar passages.

By Zhexi Feng, Ruiyi Zhang, Yongbo Yang, Pengtao Xie
arXiv AI
Sep 10

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

The paper introduces MERIT, a benchmark that evaluates the marginal benefit of long‑term memory for tool‑using large language model agents while explicitly accounting for cost. MERIT provides episodic tool‑use tasks across three domains, verifies dependence on earlier‑episode facts, and measures memory operations in tokens and dollars. Experiments on GPT‑4.1‑mini, Claude Haiku 4.5, and Claude Sonnet 5 show that memory can significantly improve task success, but its utility varies widely across models and memory implementations, and full replay is rarely cost‑effective.

By Shweta Mishra, Shashank Mishra
arXiv Machine Learning
Sep 22

EAVer: Long-Form Factuality Verification as an End-to-End Agentic Policy

arXiv:2609.22223v1 Announce Type: cross Abstract: Long-form factuality verification is commonly implemented as a static decompose-search-verify pipeline, with separately prompted modules processing c...

By Kening Zheng, Aoying Zheng, Zhigang Chang, Yazhi Guo, Miaotian Guo, Qingwei Zong, Xianhai Xie, Weiqiang Jin, Chengze Li, Hanrong Zhang, Jie Yang, Wei-Chieh Huang, Lingzhe Zhang, Liancheng Fang, Xin Zou, Hanqian Li, Jiahao Huo, Yibo Yan, Zizhuang Deng, Lei Miao, Wei Guo, Haihong Tang, Bo Zheng, Philip S. Yu