arXiv AI

The Default Trap: Rethinking Plan Evaluation in Tool-Using LLM Agents

The paper investigates how large language model agents that use tools respond to changes in plan priorities versus default plan removal, a phenomenon termed the "default trap." Experiments across 3,200 decision windows on Retail, Airline, and AgentDojo tasks show that switching priorities strongly redirects model choices, while removing a default plan yields weaker responsiveness. Additional studies reveal that the order of account lists and the presence of extra text significantly influence default target selection and priority effects, with overall task success varying from -19.4 to +8.3 points relative to no plan.

arXiv AI
Sep 18

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

The paper investigates how agent harnesses—specifically planning guidance, execution organization, and completion verification—affect performance in retail and airline pilot tasks. By comparing fixed, task‑specific plans to shuffled policy text of equal length, the study finds that fixed plans improve success rates by about 7 percentage points, especially on complex tasks. A read‑only verifier rejects a majority of invalid episodes while incurring minimal cost, and its impact varies with the penalty for erroneous acceptance, often matching the full planning‑plus‑verification benefit at a lower cost.

By Yukun Zhang, Kemu Xu, Yishen Chen
arXiv AI
3d ago

PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents

PriceBench is a diagnostic benchmark that extracts price, quality, and brand preferences from large language models (LLMs) by analyzing their hotel booking choices. Using a logit choice model, the study evaluated 28 LLMs from eight providers across 3,600 booking tasks involving 179 New York City hotels. Results show that more capable LLMs exhibit stronger, more consistent preferences, while weaker models either lock onto a single position or show near-indifference, with significant variation in price sensitivity and price/quality trade-offs across providers.

By Pavel Kireyev
arXiv AI
22h ago

What Can Component-Replacement Evidence Establish? A Critical Scoping Review of Local Decisions in LLM Agents

The article reviews how replacing components in language‑model agents affects execution trajectories and downstream outcomes. It maps 348 studies, analyzes 90 comparison records, and finds that while many studies report both local decision metrics and task endpoints, they rarely demonstrate matched comparisons or prove that improved local decisions drive task‑level gains. The review identifies three potential mechanisms—recovery and disruption, intervention timing, and downstream use—and proposes eight claim‑specific reporting items to clarify evidence quality.

By Shuyang Zhang (The Hong Kong Polytechnic University), Jianshuo Chang (The Hong Kong Polytechnic University)
arXiv Computation and Language
Sep 16

Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents

The paper introduces ATRBench, a benchmark that measures the proactivity gap in long‑lived LLM agents by evaluating their ability to ask for user preferences that are not needed immediately but may be useful in future sessions. It defines the Ask‑to‑Remember (ATR) task, where agents must decide whether to request a reusable preference now, and shows that current state‑of‑the‑art agents perform significantly below an oracle. The study identifies preference acquisition as the main bottleneck and provides a diagnostic framework for improving agent proactivity.

By Bin Wu, Guanyun Zou, Bingbing Wang, Huan Zhao, Chuan Shi