The paper investigates how agent harnesses—specifically planning guidance, execution organization, and completion verification—affect performance in retail and airline pilot tasks. By comparing fixed, task‑specific plans to shuffled policy text of equal length, the study finds that fixed plans improve success rates by about 7 percentage points, especially on complex tasks. A read‑only verifier rejects a majority of invalid episodes while incurring minimal cost, and its impact varies with the penalty for erroneous acceptance, often matching the full planning‑plus‑verification benefit at a lower cost.
By Yukun Zhang, Kemu Xu, Yishen Chen
arXiv:2608. 05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic.
By Jie Wu, Ming Gong, Feixiang Cheng, Qinqin Zhao
arXiv:2604. 12147v3 Announce Type: replace-cross Abstract: Agents are commonly instructed to follow a task-specific plan for guidance.
By Shuyang Liu, Saman Dehghan, Jatin Ganhotra, Martin Hirzel, Reyhaneh Jabbarvand
Background. A component replacement in a language-model agent changes an execution trajectory, potentially altering later observations, resource use, and recovery opportunities. Different evidence is...
arXiv:2606. 19744v1 Announce Type: cross Abstract: Aligning language models with human preferences often requires optimising multiple behavioural objectives.
By Pranav Bhandari, Nicolas Fay, Amitava Datta, Usman Naseem, Mehwish Nasim
arXiv:2609.38108v1 Announce Type: new
Abstract: Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successfu...
By Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera, Marcos L\'opez de Prado, Shadab Khan
PriceBench is a diagnostic benchmark that extracts price, quality, and brand preferences from large language models (LLMs) by analyzing their hotel booking choices. Using a logit choice model, the study evaluated 28 LLMs from eight providers across 3,600 booking tasks involving 179 New York City hotels. Results show that more capable LLMs exhibit stronger, more consistent preferences, while weaker models either lock onto a single position or show near-indifference, with significant variation in price sensitivity and price/quality trade-offs across providers.
By Pavel Kireyev
The article reviews how replacing components in language‑model agents affects execution trajectories and downstream outcomes. It maps 348 studies, analyzes 90 comparison records, and finds that while many studies report both local decision metrics and task endpoints, they rarely demonstrate matched comparisons or prove that improved local decisions drive task‑level gains. The review identifies three potential mechanisms—recovery and disruption, intervention timing, and downstream use—and proposes eight claim‑specific reporting items to clarify evidence quality.
By Shuyang Zhang (The Hong Kong Polytechnic University), Jianshuo Chang (The Hong Kong Polytechnic University)
arXiv:2607. 12216v1 Announce Type: cross Abstract: Multi-agent and memory-augmented LLM systems often place coordination content, shared state, prior discussion, tool outputs, summaries, and role instructions, inside the same finite prompt used for the current task.
By Brenda Lelis, Rodrigo Cabral-Carvalho
arXiv:2607. 07097v1 Announce Type: new Abstract: Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect.
By Lifei Liu, Haoran Yu, Xiaochong Jiang, Su Wang, Pin Qian, Yihang Chen
The paper introduces ATRBench, a benchmark that measures the proactivity gap in long‑lived LLM agents by evaluating their ability to ask for user preferences that are not needed immediately but may be useful in future sessions. It defines the Ask‑to‑Remember (ATR) task, where agents must decide whether to request a reusable preference now, and shows that current state‑of‑the‑art agents perform significantly below an oracle. The study identifies preference acquisition as the main bottleneck and provides a diagnostic framework for improving agent proactivity.
By Bin Wu, Guanyun Zou, Bingbing Wang, Huan Zhao, Chuan Shi
arXiv:2606. 23937v1 Announce Type: cross Abstract: Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model.
By Tianyu Ding, Juan Pablo De la Cruz Weinstein