FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement
arXiv:2607. 01111v1 Announce Type: cross Abstract: Robot policies inevitably encounter failures when deployed in real environments.
arXiv:2608. 11977v1 Announce Type: new Abstract: Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed reliably, yet deployed tools can fail transiently, persistently, or silently.
arXiv:2607. 01111v1 Announce Type: cross Abstract: Robot policies inevitably encounter failures when deployed in real environments.
arXiv:2608. 05080v1 Announce Type: new Abstract: Critic-free group-based reinforcement learning has become a scalable approach for post-training large language models.
The paper introduces Growing Harness, a training method that transforms recurring control logic in large language model agents into reusable executable code, reducing reliance on the model for task-specific decisions. By using strategy-free scaffolds, failure-guided code repair, and success-first gating, the approach learns a shared harness that improves performance across multiple benchmarks and model sizes. Experiments on BrowseComp-Plus and WebArena-Verified show significant gains in success rates and substantial reductions in LLM calls and inference cost compared to traditional tool‑calling agents.
The paper introduces Boundary-Aware Skill Memory (BASM), a method that enriches skill memories for large language model agents with explicit boundary fields such as applicability conditions, risk cues, avoidance rules, and recovery notes. This approach transforms retrieved skills from unconditional templates into state‑conditioned guidance, preventing the Skill Imitation Trap where more skills lead to incorrect tool usage. Experiments on three agent benchmarks and four model scales show that BASM improves task success rates, accuracy, and reduces attack success while cutting average steps compared to memory‑free baselines.
arXiv:2608. 03403v1 Announce Type: new Abstract: The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes.
arXiv:2606. 26027v1 Announce Type: cross Abstract: Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL) methods show promise for enhancing model capabilities.
arXiv:2608.21544v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can depend on tool calls and external observations r...
arXiv:2606. 01416v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents rely on orchestration layers that coordinate planning, retrieval, tool invocation, validation, memory, and recovery.
arXiv:2512. 07287v3 Announce Type: replace-cross Abstract: As intents unfold and environments change, multi-turn agents face continuously shifting decision contexts.
arXiv:2608. 14380v1 Announce Type: new Abstract: Many real-world tasks require LLM agents to interact with their environments over long execution horizons.
arXiv:2609.18304v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly tackle long-horizon tasks through multi-step environment interaction, yet a single erroneous action can...
ParaRecover is a new process-level benchmark designed to evaluate error localization and recovery in multi-turn parallel tool-use agents. It contains 10,626 instances across two difficulty levels, built on a taxonomy of 14 error types that cover planning dependencies, tool selection, and argument matching. The benchmark introduces the SDE rubric, which assesses structural integrity, diagnostic reasoning, and evolutionary strategy during agent execution, and demonstrates that it can guide improvements in agents’ reflective recovery capabilities.