arXiv:2608. 05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic.
By Jie Wu, Ming Gong, Feixiang Cheng, Qinqin Zhao
arXiv:2604. 12147v3 Announce Type: replace-cross Abstract: Agents are commonly instructed to follow a task-specific plan for guidance.
By Shuyang Liu, Saman Dehghan, Jatin Ganhotra, Martin Hirzel, Reyhaneh Jabbarvand
The paper investigates how large language model agents that use tools respond to changes in plan priorities versus default plan removal, a phenomenon termed the "default trap." Experiments across 3,200 decision windows on Retail, Airline, and AgentDojo tasks show that switching priorities strongly redirects model choices, while removing a default plan yields weaker responsiveness. Additional studies reveal that the order of account lists and the presence of extra text significantly influence default target selection and priority effects, with overall task success varying from -19.4 to +8.3 points relative to no plan.
By Xueqi Li, Jingjie Ning, Yibo Kong
arXiv:2609.38108v1 Announce Type: new
Abstract: Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successfu...
By Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera, Marcos L\'opez de Prado, Shadab Khan
Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable. We show that failure is predictable early from the agent's internal representations: lightweight per-round probes on hidden activations anticipate eventual episode failure as early as the first interaction round, where scorers reading only the agent's observable behavior are barely better than chance.
arXiv:2609.37125v1 Announce Type: new
Abstract: Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition c...
By Zhengkun Di, Bin Shi, Kai Sun, Yiming Xu, Bo Dong