arXiv:2609.17419v1 Announce Type: new
Abstract: Long-horizon LLM agents must maintain task state across extended sequences of observations, actions, tool calls, and intermediate beliefs. We study the...
By Xinyuan Song, Zekun Cai
arXiv:2609.24744v1 Announce Type: new
Abstract: Language agents solve complex tasks through plans and actions. A single step the world refuses puts the goal out of reach, and what the agent does next...
By Sungheon Jeong, Sanggeon Yun, Ryozo Masukawa, Haleh Alimohamadi, Mahdi Imani, Mohsen Imani
arXiv:2510.15047v2 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) as agents often fail to improve in new environments. We identify and characterize a failure mode we call explora...
By Shiqi Chen, Tongyao Zhu, Zian Wang, Jinghan Zhang, Kangrui Wang, Ruochen Zhou, Siyang Gao, Teng Xiao, Yee Whye Teh, Junxian He, Manling Li
LLM‑BabyBench transforms the BabyAI gridworld into a fully observable, purely textual setting that isolates planning as the sole source of failure. By serialising the entire grid, providing formal instructions, and validating actions deterministically, the benchmark introduces the PPD suite—Predict, Plan, and Decompose tasks—each scored with metrics that separate mission understanding from sequencing. Across a range of large language models, simulation accuracy is high while planning success drops sharply beyond a model‑specific horizon, revealing that plan length—not grid size—drives failure and that models often commit to a single corridor‑shaped route without backtracking.
By Idriss Malek, Omar Choukrani, Daniil Orel, Anh Duy Le Dinh, Zhuohan Xie, Zangir Iklassov, Martin Tak\'a\v{c}, Salem Lahlou
arXiv:2607. 16204v1 Announce Type: new Abstract: Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments.
By Darshan Deshpande
arXiv:2606. 24842v1 Announce Type: new Abstract: In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces.
By Yikai Lu, Yifei Wu, Xinyu Lu, Tongxin Li
arXiv:2606. 27806v3 Announce Type: replace Abstract: Language agents plan by generating not only actions but also implicit predictions of how the world will change.
By Xinyuan Song, Zekun Cai
arXiv:2608.22421v1 Announce Type: new
Abstract: World models predict action-conditioned futures and serve as critical internal simulators for downstream planning and control. However, catastrophic pr...
By Zhanpeng Shi, Zi Liang, Rong Feng, Shiqin Tang, Xuyang Chen, Hongzong Li
The paper introduces VHD-Play, a pipeline that first samples and solves a mathematical model before generating agentic reinforcement learning environments, ensuring that dynamics and evaluation are aligned from the outset. This approach yields 3,300 diverse environments at a low cost and significantly improves the performance of a large language‑model agent (Qwen3.6‑35B‑A3B) across multiple diagnostic families and external benchmarks. The study demonstrates that stateful interaction is a key factor in learning gains and that scaling the training substrate can further enhance performance.
By Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
The paper introduces Long-Transduction, a diagnostic framework designed to evaluate how well language models can maintain task fidelity during extended generation tasks that involve continuous reading, mutating, and outputting of context-dependent operations such as arithmetic, sorting, variable lookups, and table transformations. By independently varying local task complexity, input data formatting, and context length, the study isolates failure modes across these axes. Experiments on seven open-weight models reveal significant performance drops—62.8% when scaling context length from 4 to 128K, 36.5% with input format changes, and 39.9% with increased local task complexity—highlighting critical vulnerabilities in long-horizon agentic workflows.
By Jeffrey Willette, Krishna C. Puvvada, Boris Ginsburg
Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable...
arXiv:2606. 27806v1 Announce Type: new Abstract: World models for language agents come in two useful forms.
By Xinyuan Song, Zekun Cai