The study evaluates how large language model agents maintain consistency over extended interactions by simulating a 20‑step delayed‑gratification task. Researchers ran 84,540 trajectories across eight model families, using survival analysis to track when agents first claim a reward and discrete‑time hazard regression to assess how factors like social visibility, persona stressors, and deliberation policy affect failure risk. They also developed a seven‑category taxonomy from 13,780 deliberation traces, revealing that early failures are impulse‑driven, later ones are fatigue‑ or cost‑benefit‑framed, and public settings elicit norm‑oriented justifications; longer deliberation correlates with higher intra‑rationale contradictions, challenging assumptions about reasoning depth and consistency.
By Igor Bogdanov, Olga Manakina, Chung-Horng Lung
ReLiveGym is a diagnostic environment that evaluates long‑lived language‑model agents over weeks of chronologically replayed real‑world streams such as news, market data, and social media. The tasks vary in time sensitivity, reasoning depth, and recurrence, and the study tests eight base language models to see how model choice and harness design—especially action timing—affect performance. Continuous learning from hindsight feedback is also examined to address failure modes in these long‑term tasks.
By Xisen Jin, Jingheng Li, Zhenglun Chen, Junyi Du, Xiang Ren
arXiv:2609.15309v1 Announce Type: new
Abstract: Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to s...
By Kaiyuan Liu, Qiuyang Mang, Bo Peng, Wenhao Chai, Hanchen Li, Shreyas Pimpalgaonkar, Luke Zettlemoyer, Alex Dimakis, Alvin Cheung
CivBench is an open‑source benchmark that evaluates language‑model agents in the long‑horizon, tool‑mediated game Civilization VI using the Model Context Protocol (MCP). Each episode lasts over 300 turns, generating thousands of tool calls across a 76‑tool action space, and includes a narration layer that translates visual game state into structured text. The study characterises agent behaviour across four model families, introducing Proactive Monitoring Rate (PMR) and RAG@10 as interface‑level metrics, and finds that agents often under‑monitor strategic state and fail to execute near‑term commitments despite tool access and explicit guidance.
By Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster, Rui Ponte Costa
arXiv:2606. 28733v1 Announce Type: new Abstract: LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete user goals.
By Han Luo, Bingbing Wen, Lucy Lu Wang
The paper investigates why large language model (LLM) agents fail on long, multi‑step production workflows despite high benchmark success. By testing nine models (1.2 B–671 B parameters) across six task families and multiple horizons, the authors find that task success follows a geometric decay governed by a per‑step reliability that never reaches 1, leading to inevitable collapse for long horizons. The degradation is driven mainly by step count rather than context length, and the study quantifies a significant gap between benchmark and production performance, especially for agentic tool‑use tasks.
By Shubhra Mittal