arXiv AI By Xinyuan Song, Zekun Cai

When Do Causal World Models Help Modular LLM Agents

Read the original on arXiv AI →

The paper introduces FedCausalCompose, a causal world‑model framework designed for modular large‑language‑model agents that interact with distinct services such as order, payment, inventory, and shipment. It demonstrates that standard observational world models suffer from irreducible interventional errors when unblocked back‑door paths exist, whereas incorporating intervention‑response evidence improves interface recovery and can outperform non‑causal baselines when coverage and local mechanism errors are controlled. Experiments show that causal interfaces are most beneficial in structured tool environments with clear API signatures, while they provide little advantage in dialogue or narrative settings unless the causal information becomes directly relevant to the agent’s decision making.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 25

OpenRCA 2.0: From Outcome Labels to Causal Process Supervision

Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and tool use. However, existing datasets suffer from a fundamental gap: they label only the root cause, not the propagation path connecting it to the observed symptom, which largely simplifies the task to naive pattern matching.

arXiv AI
2d ago

When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents

Large language model agents depend on external harnesses to exchange information with their environment and to recover from execution errors, but recovery is typically evaluated only by overall task success, masking a key trade‑off. The authors treat recovery as a causal decision problem, comparing outcomes with and without recovery from the same execution state to separate rescue from harm and analyze how its value evolves over time. They propose the Causal Intervention Router (CIR), a lightweight policy that uses pre‑recovery information to decide when intervention is beneficial, achieving a 3‑point increase in success on long‑horizon ALFWorld tasks with Qwen3‑14B while preserving correct observations and demonstrating that recovery’s benefit is not solely due to new observations.

By Shuyao Xiao, Shengling Wang, Xuan Chen, Ke Chao, Ming Cui, Feifei Qian, Chaoyang Mei, Fanlin Meng, Ziming Yu, Junxi Yin