arXiv AI

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

CAFE (Coupled Agent–Feedback Evolution) is a framework that lets a shared‑parameter model alternate between acting as a search agent and as a critic that provides corrective feedback. By learning when to request feedback and how to use it, CAFE trains the agent to recover from its own failures and shapes rewards both online and offline. Experiments on seven search benchmarks show that CAFE outperforms other RL‑based agents, maintains gains on out‑of‑domain tests, and reduces hallucinations, indicating that co‑evolving feedback is essential for self‑improving search agents.

arXiv AI
Jul 7

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry

arXiv:2607. 03702v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effectively leverage failed trajectories: full retries incur high interaction costs, while experience retrieval tends to dilute critical experience signals.

By Weiyang Guo, Zesheng Shi, Longhui Zhang, Zeen Zhu, Min Zhang, Jing Li