arXiv AI By Xiaofang Yang, Lijun Li, Heng Zhou, Tong Zhu, Xiaoye Qu, Yuchen Fan, Qianshan Wei, Rui Ye, Li Kang, Yiran Qin, Daizong Liu, Qi Li, Ning Ding, Siheng Chen, Jing Shao

Toward Efficient Agents: Memory, Tool learning, and Planning

Read the original on arXiv AI →

arXiv:2601. 14192v2 Announce Type: replace Abstract: Recent years have witnessed increasing interest in extending large language models into agentic systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 10

Rethinking the Evaluation of Efficiency Methods for Multi-Agent Systems

The paper critiques current evaluations of efficiency methods for large language model–based multi‑agent systems, arguing that reported gains are often inflated by method‑specific prompts and starting topologies. It introduces a controlled, MAS‑demanding diagnostic benchmark that standardizes the backbone model, agent registry, and runtime, and systematically varies topology, scale, depth, and tool use. The authors find that many claimed efficiency improvements are setup‑dependent, sometimes stemming from structural collapse or random pruning rather than genuine, robust gains.

By Jiamu Zhang, Lingxi Zhang, Pengjun Lu, Qiyue Zhang, Yu-Neng Chuang, Zhengchen Li, Shuai Xu, Vipin Chaudhary, Hanjie Chen
Hugging Face Trending Papers
Aug 3

CRISP: Critical Step Perception for Training Efficient Deep Search Agents

Large language models (LLMs) are increasingly extended into deep search agents that solve complex questions through multi-step interaction with external search and browsing tools. However, existing agents often incur substantial computational and interaction costs, generating lengthy trajectories that contain redundant queries, inefficient exploration, and irrelevant observations.

arXiv Computation and Language
Sep 21

MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks

MemoryArena is a new evaluation gym that benchmarks agent memory in interdependent multi‑session tasks. Unlike prior benchmarks that test memorization or single‑session action in isolation, MemoryArena requires agents to acquire memory while interacting with the environment and then use that memory to guide future decisions across a range of tasks such as web navigation, planning, information search, and formal reasoning. The benchmark reveals that agents excelling on existing long‑context memory tests perform poorly here, highlighting a gap in current memory evaluation methods.

By Zexue He, Yu Wang, Churan Zhi, Yuanzhe Hu, Tzu-Ping Chen, Lang Yin, Ze Chen, Tong Arthur Wu, Siru Ouyang, Zihan Wang, Jiaxin Pei, Julian McAuley, Yejin Choi, Alex Pentland