arXiv AI By Igor Bogdanov, Olga Manakina, Chung-Horng Lung

Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy

Read the original on arXiv AI →

The study evaluates how large language model agents maintain consistency over extended interactions by simulating a 20‑step delayed‑gratification task. Researchers ran 84,540 trajectories across eight model families, using survival analysis to track when agents first claim a reward and discrete‑time hazard regression to assess how factors like social visibility, persona stressors, and deliberation policy affect failure risk. They also developed a seven‑category taxonomy from 13,780 deliberation traces, revealing that early failures are impulse‑driven, later ones are fatigue‑ or cost‑benefit‑framed, and public settings elicit norm‑oriented justifications; longer deliberation correlates with higher intra‑rationale contradictions, challenging assumptions about reasoning depth and consistency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 25

Delay-of-Gratification as a Multi-Agent Survival Micro-benchmark for Long-Horizon LLMs: Social Exposure, Personas, and Tool Use Budgets

The paper introduces a new multi‑agent micro‑benchmark called Delay‑of‑Gratification, modeled after the Stanford marshmallow experiment, to evaluate large language models (LLMs) in long‑horizon, multi‑turn interactions. In the benchmark, ReAct agents use a per‑step “raise a question” tool under various constraints—social context (broadcast vs. isolated), persona traits (age, hedonic drive), and tool‑use policy (mandatory vs. optional). Across 19,200 trajectories, the study finds that most agents exhibit an early impulse to “eat,” only 75.9% persist to the end, and factors such as isolation and hedonic drive significantly influence survival and questioning behavior, with ablations showing that removing hedonic drive and age can improve completion rates.

By Olga Manakina, Igor Bogdanov, Chung-Horng Lung
Hugging Face Trending Papers
Jul 7

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade

Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable. We show that failure is predictable early from the agent's internal representations: lightweight per-round probes on hidden activations anticipate eventual episode failure as early as the first interaction round, where scorers reading only the agent's observable behavior are barely better than chance.

arXiv AI
2d ago

ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality

ReLiveGym is a diagnostic environment that evaluates long‑lived language‑model agents over weeks of chronologically replayed real‑world streams such as news, market data, and social media. The tasks vary in time sensitivity, reasoning depth, and recurrence, and the study tests eight base language models to see how model choice and harness design—especially action timing—affect performance. Continuous learning from hindsight feedback is also examined to address failure modes in these long‑term tasks.

By Xisen Jin, Jingheng Li, Zhenglun Chen, Junyi Du, Xiang Ren
arXiv AI
Sep 3

How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making

The paper investigates why large language model (LLM) agents fail on long, multi‑step production workflows despite high benchmark success. By testing nine models (1.2 B–671 B parameters) across six task families and multiple horizons, the authors find that task success follows a geometric decay governed by a per‑step reliability that never reaches 1, leading to inevitable collapse for long horizons. The degradation is driven mainly by step count rather than context length, and the study quantifies a significant gap between benchmark and production performance, especially for agentic tool‑use tasks.

By Shubhra Mittal