arXiv:2607. 03177v1 Announce Type: cross Abstract: Traditional reinforcement learning (RL) for recovery in autonomous systems lacks causal understanding and generalizes poorly to novel failure scenarios.
By Safia Fatima, Kai Olav Ellefsen, Leon Moonen
The paper introduces FedCausalCompose, a causal world‑model framework designed for modular large‑language‑model agents that interact with distinct services such as order, payment, inventory, and shipment. It demonstrates that standard observational world models suffer from irreducible interventional errors when unblocked back‑door paths exist, whereas incorporating intervention‑response evidence improves interface recovery and can outperform non‑causal baselines when coverage and local mechanism errors are controlled. Experiments show that causal interfaces are most beneficial in structured tool environments with clear API signatures, while they provide little advantage in dialogue or narrative settings unless the causal information becomes directly relevant to the agent’s decision making.
By Xinyuan Song, Zekun Cai
Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. Failed runs tend to be longer and exhibit redundant exploration or looping, suggesting that some failures may be detectable before completion.
arXiv:2608. 03222v1 Announce Type: cross Abstract: Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates.
By Chenyu Wang, Yunbo Lyu, Junda He, Zhou Yang, Chenxing Zhong, Yaniv Harel, David Lo
arXiv:2606. 08275v1 Announce Type: cross Abstract: When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observability) or whether it passed (evaluation), but not which step caused the failure.
By Jaineet Shah
Self-correction is particularly useful when a failure constrains the next repair. Coding agents benefit from this property because compilers, tests, and execution traces turn many failures into typed recovery signals, but broad language-agent tasks often expose only a coarse task failure.
The paper introduces the concept of causal retention in interactive agents, examining whether a frozen learned state can correctly answer a mechanism‑probe map that is fixed independently of training. It shows that for finite structural causal models the optimal probe error is a Bayes decision risk, vanishing only when each learning‑interface fiber lies within a single probe‑answer fiber, and provides theoretical results such as a posterior‑coverage theorem and an exact edit decomposition. Experiments on finite causal systems, continuous simulators, TD‑MPC2, and Qwen2.5‑7B‑Instruct demonstrate that causal retention can be achieved with high accuracy, outperforming task‑performance‑based approaches.
By Shengjun Zhang, Tingyi Liu, Dong Xie, Yunlong Dong, Xiang Wang, Cheng Zeng
The paper introduces a new evaluation protocol called checkpoint handoff to disentangle the contributions of reaching a target state and solving the task in reinforcement learning agents. By cloning states reached by one checkpoint and handing them to another without retraining, the authors separate the REACH metric (how often a policy arrives at a state confirmed to be a fixed number of actions from success) from the SOLVE metric (how often it finishes from that identical state). Across two benchmarks and pipelines, the analysis shows that RL history benefits RL solvers more than SFT solvers, and that independent REACH and SOLVE gaps predict overall performance.
By Xuan Liu, Jingbin Qian
arXiv:2609.13672v1 Announce Type: new
Abstract: AI agents can be interrupted while editing files, calling tools, or carrying out multi-step tasks. Restarting repeats completed work, but continuing fr...
By Zhihui Zhang, Wei Liu
arXiv:2608.25920v2 Announce Type: replace
Abstract: As large language model (LLM)-based multi-agent systems (MASs) are increasingly applied to long-horizon complex tasks, their reliability has emerge...
By Zhongwen Luan, Xiaoyu Zhang, Ming Hu, Yue Yang, Jiongchi Yu, Xiaohong Chen
arXiv:2609.18304v1 Announce Type: new
Abstract: Large language model (LLM) agents increasingly tackle long-horizon tasks through multi-step environment interaction, yet a single erroneous action can...
By Yi Yu, Liuyi Yao, Yaliang Li, Enshu Wang, Libing Wu
arXiv:2606. 16330v1 Announce Type: new Abstract: Disruption recovery in industrial assembly lines requires timely decisions under machine faults, worker absence, and emergency orders.
By Xin Huang, Yongcai Wang, Fengyi Zhang, Zhikun Tao, Yunjun Han, Naiqi Wu