The paper introduces the Multi-Session Personalized Tool Calling (MPT) benchmark, containing 4,695 instances across 459 multi‑session histories that test Preference Recall, Induction, and Transfer. It proposes PRefine, a test‑time memory method that refines a user’s latent preference via a generate‑verify‑refine loop. Experiments with five LLMs show that PRefine outperforms existing memory systems and even full‑history prompting on Preference Transfer, suggesting that personalized agents should encode behavior as preferences rather than merely storing past interactions.
By Yejin Yoon, Minseo Kim, Taeuk Kim
arXiv:2607. 12385v1 Announce Type: new Abstract: A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing.
By Genglin Liu, Saadia Gabriel
S3Gym is an interactive benchmark designed to evaluate large language models (LLMs) on their ability to self-improve through self-testing, self-judging, and self-improvement. It separates permissive exploration from strict held-out evaluation across seven text-based games with executable environment verifiers. Experiments show that self-improvement varies by task, with different experience incorporation pathways (direct history, summary memory, or parameter training) yielding mixed results and highlighting the need for agents to transform feedback into executable, transferable policies.
By Jiajun Shi, Siyuan Tao, Yuhao Wu, Zexuan Wang, Jingyuan Zhang, Jiaheng Liu, Xinping Lei, Xinrong Zhang, Siyuan Fang, Zhewen Tan, Tianle Cai, Junhao Fang, Jiameng Huang, Yueyang Wang, Jinkai Liu, Yuxuan Zhang, Jian Yang, Zhoujun Li, Shen Yan, Wenhao Huang, Ge Zhang
arXiv:2609.38043v1 Announce Type: new
Abstract: Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated u...
By Ashish Jain, Armaan Sandhu
arXiv:2606. 02965v2 Announce Type: replace Abstract: As large language models gain tool access and are deployed as autonomous agents capable of editing records, executing transactions, and modifying infrastructure, we still evaluate them based on the sole metric of task completion.
By Victor Ojewale, Suresh Venkatasubramanian
The paper investigates how personalized agents decide to use, ignore, update, or query retrieved user memory before acting on a task. An empirical audit protocol is developed to test structured intermediate outputs, revealing that while exposing state definitions improves accuracy, an explicit state-output field does not significantly enhance policy accuracy for large language models. The study also shows that example-level accuracy overstates consistency, with full four‑way family success being rare, and that providing benchmark‑associated state labels merely conditions predictions rather than proving internal fidelity.
By Yihang Chen, Pin Qian, Su Wang, Chong Peng, Huan Xu, Shuaiting Li, Yiqi Sun
LLM-based agents have shown strong capabilities in automated data analysis and are increasingly moving toward long-horizon, multi-stage analytical workflows. However, as the analytical process evolves...
arXiv:2609. 04875v1 Announce Type: cross Abstract: Long-running LLM agents are stateful: beyond the transcript they accrete compressed summaries, plaintext memory, pending tool plans, and, under every serving API, a KV cache.
By Chao Yao, Yangbo Wei, Zhen Huang, Junhong Qian, Chenle Chen, Shaoqiang Lu, Chen Wu, Lei He
arXiv:2606. 02965v1 Announce Type: new Abstract: Benchmarks for autonomous agents measure whether agents complete tasks, yet this framing is systematically blind to whether an agent should have proceeded at all.
By Victor Ojewale, Suresh Venkatasubramanian
arXiv:2606. 05922v1 Announce Type: cross Abstract: AI agents rely on a harness of skills, tools, and workflows to solve complex problems.
By Wenbo Pan, Shujie Liu, Chin-Yew Lin, Jingying Zeng, Xianfeng Tang, Xiangyang Zhou, Yan Lu, Xiaohua Jia
AI agents rely on a harness of skills, tools, and workflows to solve complex problems. Continually improving this harness is essential for adapting to new tasks.
MemoryArena is a new evaluation gym that benchmarks agent memory in interdependent multi‑session tasks. Unlike prior benchmarks that test memorization or single‑session action in isolation, MemoryArena requires agents to acquire memory while interacting with the environment and then use that memory to guide future decisions across a range of tasks such as web navigation, planning, information search, and formal reasoning. The benchmark reveals that agents excelling on existing long‑context memory tests perform poorly here, highlighting a gap in current memory evaluation methods.
By Zexue He, Yu Wang, Churan Zhi, Yuanzhe Hu, Tzu-Ping Chen, Lang Yin, Ze Chen, Tong Arthur Wu, Siru Ouyang, Zihan Wang, Jiaxin Pei, Julian McAuley, Yejin Choi, Alex Pentland