arXiv AI

Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)

arXiv:2605. 09169v2 Announce Type: replace-cross Abstract: A Mamba state-space model trained only for next-step prediction appears to recover Granger-causal structure through a simple readout $S = |W_{out} W_{in}|$, with early experiments suggesting the phenomenon generalized across architectures and benefited from interventional data at $p < 10^{-5}$.

arXiv Machine Learning
Jul 31

DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series

arXiv:2607. 27263v1 Announce Type: new Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation under-served exactly where it matters most, such as in healthcare, policy evaluation, and climate science.

By Dennis Thumm, Billy Tim Anthony, Ying Chen
arXiv Machine Learning
Sep 4

Guide, Not Bind: Why Defeasible Priors Fail in Augmented Lagrangian Causal Discovery

The paper investigates why differentiable causal discovery methods that encode expert priors as forbidden-edge constraints via an Augmented Lagrangian (ALM) penalty—termed the "guide, not bind" approach—often fail. It identifies two key failures: (1) the sequential penalty‑ramping ALM suppresses a true edge before counterfactual checks can detect it, and the proposed adaptive relaxation rule DADU violates necessary conditions for safe relaxation, leading to a high failure rate across thousands of training runs; (2) the standard correlation‑matching objective inherently ties a true edge and its reverse to the same cost, whereas covariance matching can separate them by a provable margin. The authors provide theoretical propositions, corollaries, and empirical evidence to support these claims.

By Sairam Sundararaman, Sara Girdhar, Manit Narasimha Murthy, Samrudh N, Bhaskarjyoti Das
arXiv AI
Sep 24

Are Stated Reasoning Steps Causally Load-Bearing?

The study investigates whether the reasoning steps a language model writes are causally responsible for its answers. Using a causal intervention method on the activation stream, the authors find that for Qwen3-4B, about 77% of stated steps are causally load‑bearing, while behavioral tests overestimate this by roughly 11 percentage points. The faithfulness of reasoning decreases with model size and depth of reasoning, especially for the smaller Qwen3-1.7B.

By Abhiram Bhupatiraju, Rayan Nyaupane
arXiv Machine Learning
Sep 11

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.

By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye