The paper introduces a new multi‑agent micro‑benchmark called Delay‑of‑Gratification, modeled after the Stanford marshmallow experiment, to evaluate large language models (LLMs) in long‑horizon, multi‑turn interactions. In the benchmark, ReAct agents use a per‑step “raise a question” tool under various constraints—social context (broadcast vs. isolated), persona traits (age, hedonic drive), and tool‑use policy (mandatory vs. optional). Across 19,200 trajectories, the study finds that most agents exhibit an early impulse to “eat,” only 75.9% persist to the end, and factors such as isolation and hedonic drive significantly influence survival and questioning behavior, with ablations showing that removing hedonic drive and age can improve completion rates.
By Olga Manakina, Igor Bogdanov, Chung-Horng Lung
Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable. We show that failure is predictable early from the agent's internal representations: lightweight per-round probes on hidden activations anticipate eventual episode failure as early as the first interaction round, where scorers reading only the agent's observable behavior are barely better than chance.
ReLiveGym is a diagnostic environment that evaluates long‑lived language‑model agents over weeks of chronologically replayed real‑world streams such as news, market data, and social media. The tasks vary in time sensitivity, reasoning depth, and recurrence, and the study tests eight base language models to see how model choice and harness design—especially action timing—affect performance. Continuous learning from hindsight feedback is also examined to address failure modes in these long‑term tasks.
By Xisen Jin, Jingheng Li, Zhenglun Chen, Junyi Du, Xiang Ren
The paper investigates why large language model (LLM) agents fail on long, multi‑step production workflows despite high benchmark success. By testing nine models (1.2 B–671 B parameters) across six task families and multiple horizons, the authors find that task success follows a geometric decay governed by a per‑step reliability that never reaches 1, leading to inevitable collapse for long horizons. The degradation is driven mainly by step count rather than context length, and the study quantifies a significant gap between benchmark and production performance, especially for agentic tool‑use tasks.
By Shubhra Mittal
arXiv:2608. 13598v1 Announce Type: new Abstract: Agent evaluation relies almost entirely on outcome metrics such as success rate, which capture whether an agent succeeds but not how consistently it behaves.
By Amritesh Banerjee, Pranil Raichura
arXiv:2607. 06503v1 Announce Type: new Abstract: Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable.
By Kai Ruan, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun