arXiv:2607. 27263v1 Announce Type: new Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation under-served exactly where it matters most, such as in healthcare, policy evaluation, and climate science.
By Dennis Thumm, Billy Tim Anthony, Ying Chen
arXiv:2609.31315v1 Announce Type: cross
Abstract: Unobserved common causes are pervasive in real-world time series and can induce spurious associations that causal discovery methods mistake for direc...
By Mohammad Fesanghary
arXiv:2608. 11797v1 Announce Type: new Abstract: Model merging by task arithmetic works until it doesn't, and the field diagnoses why with magnitudes: layerwise representation bias, deviations from cross-task linearity, parameter overlap.
By Chencheng Zhu
The paper investigates why differentiable causal discovery methods that encode expert priors as forbidden-edge constraints via an Augmented Lagrangian (ALM) penalty—termed the "guide, not bind" approach—often fail. It identifies two key failures: (1) the sequential penalty‑ramping ALM suppresses a true edge before counterfactual checks can detect it, and the proposed adaptive relaxation rule DADU violates necessary conditions for safe relaxation, leading to a high failure rate across thousands of training runs; (2) the standard correlation‑matching objective inherently ties a true edge and its reverse to the same cost, whereas covariance matching can separate them by a provable margin. The authors provide theoretical propositions, corollaries, and empirical evidence to support these claims.
By Sairam Sundararaman, Sara Girdhar, Manit Narasimha Murthy, Samrudh N, Bhaskarjyoti Das
arXiv:2609.06941v1 Announce Type: new
Abstract: Causal effect estimation asks how an outcome would change under an intervention, and medicine, economics, and public policy all treat it as a foundatio...
By Haohao Zhou
arXiv:2607. 25546v1 Announce Type: new Abstract: Given a model that is already trained, which features does it rely on causally versus spuriously?
By Athanasios Vlontzos, Giorgos Papanastasiou, Bernhard Kainz, Sotirios Tsaftaris
arXiv:2411.05625v2 Announce Type: replace
Abstract: We propose a new approach to falsify causal discovery algorithms without ground truth, which is based on testing the causal model on a variable pai...
By Daniela Schkoda, Philipp Faller, Patrick Bl\"obaum, Dominik Janzing
arXiv:2607. 29484v1 Announce Type: cross Abstract: Interventional data is widely regarded as the gold standard for teaching models causal reasoning.
By Xining Xun
The study investigates whether the reasoning steps a language model writes are causally responsible for its answers. Using a causal intervention method on the activation stream, the authors find that for Qwen3-4B, about 77% of stated steps are causally load‑bearing, while behavioral tests overestimate this by roughly 11 percentage points. The faithfulness of reasoning decreases with model size and depth of reasoning, especially for the smaller Qwen3-1.7B.
By Abhiram Bhupatiraju, Rayan Nyaupane
arXiv:2609.07944v1 Announce Type: new
Abstract: Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recove...
By Yonghong Zhang, Ricardo Correia, Isabel M. Parra, Yong Xie
arXiv:2607. 25532v1 Announce Type: new Abstract: Consider a model trained at a single hospital to predict patient recovery, where the measured feature $X$ bundles the patient's true health signal ($C$) with a systematic artefact from that hospital's equipment ($S$).
By Athanasios Vlontzos, Giorgos Papanastasiou, Bernhard Kainz, Sotirios Tsaftaris
CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.
By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye