arXiv AI

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

The paper investigates how agent harnesses—specifically planning guidance, execution organization, and completion verification—affect performance in retail and airline pilot tasks. By comparing fixed, task‑specific plans to shuffled policy text of equal length, the study finds that fixed plans improve success rates by about 7 percentage points, especially on complex tasks. A read‑only verifier rejects a majority of invalid episodes while incurring minimal cost, and its impact varies with the penalty for erroneous acceptance, often matching the full planning‑plus‑verification benefit at a lower cost.

arXiv AI
1d ago

The Default Trap: Rethinking Plan Evaluation in Tool-Using LLM Agents

The paper investigates how large language model agents that use tools respond to changes in plan priorities versus default plan removal, a phenomenon termed the "default trap." Experiments across 3,200 decision windows on Retail, Airline, and AgentDojo tasks show that switching priorities strongly redirects model choices, while removing a default plan yields weaker responsiveness. Additional studies reveal that the order of account lists and the presence of extra text significantly influence default target selection and priority effects, with overall task success varying from -19.4 to +8.3 points relative to no plan.

By Xueqi Li, Jingjie Ning, Yibo Kong
Hugging Face Trending Papers
Jul 7

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade

Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable. We show that failure is predictable early from the agent's internal representations: lightweight per-round probes on hidden activations anticipate eventual episode failure as early as the first interaction round, where scorers reading only the agent's observable behavior are barely better than chance.

Hugging Face Trending Papers
Jun 19

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents

Translating natural-language planning intent into verified plans is a longstanding challenge: people communicate goals in language, while classical planners require formal PDDL specifications. Recent agentic frameworks bridge this gap by orchestrating a pool of specialized repair agents inside a verifier-checked refinement loop, but the orchestrator at the centre is itself a prompted frontier LLM, paying a frontier-LLM API call at every refinement step.

arXiv AI
6d ago

ERRAND: Budgeted Maintenance of Agent Memory

ERRAND is a new method for budgeted maintenance of agent memory that treats revalidation of stored knowledge as a priced errand competing for scarce actions. It uses an errand index that is single‑peaked, allowing certainty in either direction to cost nothing, and repairs by writing new versions rather than deleting old ones. In experiments across two drifting tool‑use worlds, ERRAND outperforms non‑oracle policies, achieving up to 10.0 percentage points improvement over eager revalidation while using only 11.0% of steps, and it self‑terminates when no budget is imposed.

By Beining Wu, Zihao Ding, Jun Huang
arXiv AI
Sep 18

The Organization of Inference: Information, Resource Constraints, and AI Production

The paper investigates how the distribution of capacity and task information across stages of AI production affects the economic value of inference. Using controlled workflow experiments on software‑engineering tasks, it finds that direct execution achieves a 59.6% success rate at token ceilings of 12,000 and 24,000, while information‑constrained planning improves from 36.2% to 51.2%. The study also shows that giving planners access to task issues boosts success, and that scaling token limits changes the balance between planning and execution workloads.

By Yukun Zhang, Kemu Xu, Yishen Chen
arXiv Machine Learning
Aug 26

PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage

PinSieve is a production system that selectively serves vision‑language models (VLMs) for enterprise content‑quality triage, operating only on the grey‑zone cases that lightweight models cannot resolve. The deployed VLM Serving Agent filters 2.05× more non‑actionable items, improves review productivity by 25.7%, cuts operating costs by 16.2%, and delivers signals the same day instead of the next. A governed memory flywheel with selective feedback, audit sampling, and a bounded proposal‑verifier loop further reduces false‑negative rates from 17.73% to 13.29% over six months, while a reasoning review agent audits teacher‑generated rationales for keep/repair/drop decisions. whyItMatters":"The system demonstrates how selective VLM serving and governed feedback loops can substantially improve efficiency, cost, and accuracy in enterprise AI content‑quality pipelines."

By Chuqing Gao, Yuanfang Song, Jonathan Zhang, Yifan Wu, Vishwakarma Singh, Qinglong Zeng, Andrey Gusev