Large language model agents depend on external harnesses to exchange information with their environment and to recover from execution errors, but recovery is typically evaluated only by overall task success, masking a key trade‑off. The authors treat recovery as a causal decision problem, comparing outcomes with and without recovery from the same execution state to separate rescue from harm and analyze how its value evolves over time. They propose the Causal Intervention Router (CIR), a lightweight policy that uses pre‑recovery information to decide when intervention is beneficial, achieving a 3‑point increase in success on long‑horizon ALFWorld tasks with Qwen3‑14B while preserving correct observations and demonstrating that recovery’s benefit is not solely due to new observations.
By Shuyao Xiao, Shengling Wang, Xuan Chen, Ke Chao, Ming Cui, Feifei Qian, Chaoyang Mei, Fanlin Meng, Ziming Yu, Junxi Yin
arXiv:2606. 30537v1 Announce Type: cross Abstract: Autonomous driving policies should be able to improve continually as deployment exposes them to increasingly diverse and long-tail traffic situations.
By Cheng Gong, Haoyang Wang, Chao Lu, Zirui Li, Jianwei Gong
arXiv:2606. 16330v1 Announce Type: new Abstract: Disruption recovery in industrial assembly lines requires timely decisions under machine faults, worker absence, and emergency orders.
By Xin Huang, Yongcai Wang, Fengyi Zhang, Zhikun Tao, Yunjun Han, Naiqi Wu
MAGMA-GEN is an on‑policy data‑generation pipeline that transforms ambiguous failures in hierarchical robotic manipulation into validated recovery supervision. It uses a privileged coach to hypothesize early decision‑level errors and proposes localized corrections, then retains only those candidates that improve downstream progress when re‑executed from the same state. This approach generates supervised examples from the agent’s own failure distribution, enabling improved task success and recovery without requiring per‑step human demonstrations.
By Loan Bernat (LAAS-GEPETTO), Matthieu Grard (LAAS-RAP), Ariane Herbulot (LAAS-RAP), Florent Lamiraux (LAAS-GEPETTO)
arXiv:2607. 14826v1 Announce Type: cross Abstract: Safe physical AI for robot actions are required not only likely to succeed but tested to be safe before execution.
By Naren Vasantakumaar, Tom Schierenbeck, Michael Beetz
The paper introduces Safe Contrastive Reinforcement Learning (Safe-CRL), a method that corrects bias in contrastive RL caused by failure-terminated Markov decision processes. By applying mass-weighted InfoNCE and a log-survival-mass score, Safe-CRL uses only a one-bit failure signal to improve survival and goal-reaching performance across twelve robot navigation and locomotion tasks. The approach demonstrates complex failure-avoidance behaviors and completes the theoretical foundation of contrastive RL under failure termination.
By Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu