arXiv AI

From Training to Deployment: Post-Hoc Causal Feature Identification via Sensitivity Ratios

arXiv:2607. 25546v1 Announce Type: new Abstract: Given a model that is already trained, which features does it rely on causally versus spuriously?

arXiv AI
Jul 23

Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models

arXiv:2607. 19618v1 Announce Type: cross Abstract: Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled procedure for verifying that an apparent ``concept'' inside a model is real rather than an artifact of sequence composition.

By Sarwan Ali
arXiv AI
Jun 16

Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)

arXiv:2605. 09169v2 Announce Type: replace-cross Abstract: A Mamba state-space model trained only for next-step prediction appears to recover Granger-causal structure through a simple readout $S = |W_{out} W_{in}|$, with early experiments suggesting the phenomenon generalized across architectures and benefited from interventional data at $p < 10^{-5}$.

By Ankit Hemant Lade, Sai Krishna Jasti, Indar Kumar, Aman Chadha
arXiv AI
Aug 26

From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

The study evaluates 12 instruction‑tuned open‑weight LLMs on six causal‑graph benchmarks, testing five prompting strategies and four confidence sources. Findings show that LLMs tend to over‑predict edges, misclassify indirect or reversed edges as direct, and exhibit high over‑confidence, while conventional confidence estimates are unreliable and agreement signals offer limited improvement. The results suggest LLMs should be used as externally validated soft causal priors rather than definitive causal‑structure evidence.

By Amit Kumar, Elnur Adl Zarabi, Suranjana Trivedy, Zhiqian Chen, Lei Zhang, Kaiqun Fu, Taoran Ji
arXiv AI
Sep 17

Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

Linear probes can decode safety‑relevant concepts such as truthfulness from language‑model activations, but probe accuracy may reflect only decodability, not causal influence on model behavior. The authors show that probe weight geometry alone cannot identify the features the model actually uses, because geometrically aligned features need not be causally relevant. They introduce a sparse‑autoencoder (SAE) decomposition that ranks features by probe alignment and gradient sensitivity, and demonstrate that ablating shared, probe‑only, and random feature sets reveals a sharp dissociation: shared features drive model output changes far more than probe‑only or random features, confirming that causal relevance requires intervention beyond weight geometry.

By Devesh Tiwari, Camille Davis, Shivank Sinha, Talia Weaver, Aditya Shah, Maheep Chaudhary
arXiv Statistics ML
Aug 25

Model-Agnostic Covariate-Assisted Inference on Partially Identified Causal Effects

The paper introduces a model‑agnostic inference framework for partially identified causal effects that leverages covariate information without requiring discrete covariates or accurate conditional distribution estimates. Using duality theory for optimal transport, the method delivers uniformly valid inference in randomized experiments, is doubly robust in observational settings, achieves asymptotic unbiasedness when nuisance parameters converge semiparametrically, and allows multiplier‑bootstrap selection of covariates and models while remaining computationally efficient. Empirical applications show the approach consistently narrows identified sets and confidence intervals without imposing extra structural assumptions.

By Wenlong Ji, Lihua Lei, Asher Spector