arXiv AI
1d ago

Sentry: Learning to Recover from LLM Agent Failures at Test Time

Sentry is a failure‑management layer for large language model agents that learns from failures at test time. It retrieves relevant lessons from an external playbook when a failure occurs, verifies recovery without task rewards, and stores new lessons only if recovery succeeds, keeping the playbook out of the agent’s context. Across multiple benchmarks, Sentry outperforms both runtime‑intervention and context‑evolution baselines, and its lessons transfer to unseen tasks.

By Changxiu Ji, Amy Lu, Qizheng Zhang, Kunle Olukotun