arXiv:2510. 00615v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as agents in dynamic real-world environments, where success depends on maintaining precise records of actions and observations.
By Minki Kang, Wei-Ning Chen, Dongge Han, Huseyin A. Inan, Lukas Wutschitz, Yanzhi Chen, Robert Sim, Saravan Rajmohan
ContextPilot is a proactive context‑management framework designed to improve long‑horizon agentic reasoning with large language models. It expands the toolset to include planning, long‑term memory, and soft context offloading, and introduces a reinforcement‑learning strategy that focuses on critical editing decisions and assigns action‑level advantages. Experiments on long‑context QA and deep search tasks demonstrate that ContextPilot achieves stronger performance with a more compact working context, outperforming existing baselines across various base models and benchmarks.
By Zhuoshi Pan, Qizhi Pei, Junru Lu, Honglin Lin, H. Vicky Zhao, Di Yin, Xing Sun
arXiv:2607. 23809v1 Announce Type: new Abstract: Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment.
By Xiaochuan Li, Ryan Ming, Meng Chu, Shuai Shao, Rong Jin, Chenyan Xiong
arXiv:2606. 16432v1 Announce Type: cross Abstract: User instructions are often underspecified because humans rely on implicit assumptions about the surrounding environment.
By Lai Jiang, Cheng Qian, Zhenhailong Wang, Pan Lu, Heng Ji, Hao Peng
arXiv:2608. 06503v1 Announce Type: new Abstract: Recurrent context compression controls context growth in long-horizon agents, but its behavioral effects remain poorly understood.
By Guanghui Min, Liang Wu, Mayank Darbari, Chen Chen, Liangjie Hong
arXiv:2604.07487v2 Announce Type: replace
Abstract: Large language model agents rely on effective model context to obtain task-relevant information for decision-making. Many existing context engineer...
By Linbo Liu, Guande Wu, Han Ding, Yawei Wang, Qiang Zhou, Yuzhe Lu, Zhichao Xu, Huan Song, Panpan Xu, Lin Lee Cheong
StateComp introduces a method for long‑horizon agents to decide when to compress historical interactions based on the current agent state, rather than relying on fixed windows or periodic schedules. The framework uses a two‑stage annotation process to create KEEP and READY labels, trains an imbalance‑aware router on frozen language model representations, and groups adjacent READY interactions into compact summaries. Experiments on WorkBuddyBench show that StateComp cuts agent and summarization tokens by 52.27% and speeds up representation extraction 12.67‑fold while preserving task performance.
By Mingxuan Wang, Hongyue Chen, Yinglong Guo, Fei Luo, Chao Ning, Bo Wang, Guorun Yao, Yanbiao Ma, Jungong Han
arXiv:2609.08318v1 Announce Type: cross
Abstract: The transition from human-centric assistance to Autonomous Software Engineering (ASE) agents has enabled the resolution of complex real-world SE task...
By Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, Shikun Zhang
AnyAct introduces a universal action layer that consolidates diverse tool capabilities into a self‑evolving action space for AI agents operating in open‑world environments. It tackles the scale dilemma, tool non‑stationarity, and heterogeneous feedback by using hierarchical progressive retrieval and test‑time reliability evolution, while a heterogeneous observation grounding module unifies multi‑modal feedback. Evaluations on LiveMCPBench and the newly created OSMCP benchmark show state‑of‑the‑art performance, with significant gains in task success rate and reduced execution steps, especially for models with limited native capabilities.
By Lingrui Xu, Yangqin Jiang, Jiachang Zhang, Xubin Ren, Chao Huang
arXiv:2606. 11680v1 Announce Type: new Abstract: Large language model (LLM) agents struggle with long-horizon tasks due to their inherent statelessness, requiring all task-relevant information to be encoded in growing input contexts.
By Hao-Lun Hsu, Nikki Lijing Kuang, Boyi Liu, Zhewei Yao, Yuxiong He
The paper introduces FOCUS, a training‑free framework that compresses the interaction history of large language model agents by preserving only the past interactions that causally influence future decisions. Unlike prior methods that learn compression policies offline, FOCUS operates entirely at test time, requiring no additional data collection or fine‑tuning and can be applied to any closed‑API model. Experiments on a variety of agentic benchmarks show that FOCUS reduces peak context length by up to 48% and dependency by 73%, while improving task success by up to 8.9 percentage points.
By Shantanu Dixit, Anson Bastos, Xuchao Zhang, Chetan Bansal, Saravan Rajmohan
Large language model (LLM) agents often perform poorly on complex, long-horizon tasks because their context becomes increasingly cluttered over time. As interactions accumulate, detailed execution tra...