arXiv AI By Yukun Zhang, Kemu Xu, Yishen Chen

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

Read the original on arXiv AI →

The paper investigates how agent harnesses—specifically planning guidance, execution organization, and completion verification—affect performance in retail and airline pilot tasks. By comparing fixed, task‑specific plans to shuffled policy text of equal length, the study finds that fixed plans improve success rates by about 7 percentage points, especially on complex tasks. A read‑only verifier rejects a majority of invalid episodes while incurring minimal cost, and its impact varies with the penalty for erroneous acceptance, often matching the full planning‑plus‑verification benefit at a lower cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
1d ago

The Default Trap: Rethinking Plan Evaluation in Tool-Using LLM Agents

The paper investigates how large language model agents that use tools respond to changes in plan priorities versus default plan removal, a phenomenon termed the "default trap." Experiments across 3,200 decision windows on Retail, Airline, and AgentDojo tasks show that switching priorities strongly redirects model choices, while removing a default plan yields weaker responsiveness. Additional studies reveal that the order of account lists and the presence of extra text significantly influence default target selection and priority effects, with overall task success varying from -19.4 to +8.3 points relative to no plan.

By Xueqi Li, Jingjie Ning, Yibo Kong
Hugging Face Trending Papers
Jul 7

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade

Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable. We show that failure is predictable early from the agent's internal representations: lightweight per-round probes on hidden activations anticipate eventual episode failure as early as the first interaction round, where scorers reading only the agent's observable behavior are barely better than chance.