arXiv Machine Learning

CIDER-FM: Foundation Models for Causal Inference from Diverse Experimental Regimes

CIDER-FM is a causal foundation model that combines finite observational data with surrogate-interventional datasets to predict target conditional interventional distributions more accurately than using observational data alone. It employs an intervention-aware representation and hierarchical three‑axis attention to integrate information across variables, samples, and experimental regimes. Experiments on synthetic graphs, simulated data, and real‑world Causal Chambers data show that incorporating experimental context improves CID prediction performance.

arXiv Machine Learning
Sep 11

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.

By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
Hugging Face Trending Papers
Sep 10

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena is a unified, evolvable benchmark designed to evaluate causal discovery methods across diverse structural causal models (SCMs). It incorporates synthetic SCMs for controlled structural variation, semantic operational SCMs for human-auditable environments, and formula-grounded SCMs to test discovery under explicit scientific mechanisms, along with real-world datasets for external validity. Experiments show that performance rankings vary significantly across SCM families and protocols, indicating that strong results on one benchmark do not generalize to others, especially in the context of causal discovery foundation models.

arXiv Machine Learning
Aug 19

TabCausal: Pretraining Across Causal Environments for Tabular Causal Discovery

TabCausal is a causal discovery foundation model that learns to map datasets directly to causal graphs by pretraining across diverse causal environments. It uses a dynamic task construction strategy to expose the model to varied graph priors, mechanisms, noise models, dimensions, sample sizes, and intervention regimes, improving transferability from observational and mixed‑interventional data. On large synthetic benchmarks and a new protocol‑guided semantic benchmark, TabCausal outperforms many classical baselines and shows robust structure recovery, especially when interventional evidence is available.

By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv Machine Learning
Jun 26

Use What You Know: Causal Foundation Models with Partial Graphs

arXiv:2602. 14972v2 Announce Type: replace Abstract: Estimating causal quantities traditionally relies on bespoke estimators tailored to specific assumptions.

By Arik Reuter, Anish Dhir, Cristiana Diaconu, Jake Robertson, Ole Ossen, Frank Hutter, Adrian Weller, Mark van der Wilk, Bernhard Sch\"olkopf
arXiv AI
Jul 14

CDFM: Towards a General-Purpose Causal Discovery Foundation Model

arXiv:2607. 11508v1 Announce Type: cross Abstract: Causal discovery, the process of recovering underlying causal structures from observational data, is a fundamental pursuit across scientific disciplines.

By Jie Qiao, Ruichu Cai, Zijian Li, Weilin Chen, Pengfei Hua, Boyan Xu, Zhengming Chen, Zhifeng Hao, Peng Cui
arXiv Machine Learning
Sep 4

Causal Foundation Models

Causal Foundation Models (CFMs) are pretrained neural networks designed to estimate causal quantities—such as the average treatment effect—across new datasets using in‑context learning, eliminating the need for bespoke pipelines or model updates. The paper introduces CFMs, reviews foundational concepts in causal inference and machine learning, and provides practical code examples and Jupyter notebooks to illustrate their application.

By Christopher Stith, Hossein Rahmani, Jesse C. Cresswell