arXiv AI

A Survey on Federated Causal Discovery and Inference

arXiv:2606. 23741v1 Announce Type: cross Abstract: Causal reasoning, which encompasses the discovery of causal structures and the inference of causal effects, is fundamental to data-driven decision making.

arXiv AI
Jul 14

CDFM: Towards a General-Purpose Causal Discovery Foundation Model

arXiv:2607. 11508v1 Announce Type: cross Abstract: Causal discovery, the process of recovering underlying causal structures from observational data, is a fundamental pursuit across scientific disciplines.

By Jie Qiao, Ruichu Cai, Zijian Li, Weilin Chen, Pengfei Hua, Boyan Xu, Zhengming Chen, Zhifeng Hao, Peng Cui
Hugging Face Trending Papers
Sep 10

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena is a unified, evolvable benchmark designed to evaluate causal discovery methods across diverse structural causal models (SCMs). It incorporates synthetic SCMs for controlled structural variation, semantic operational SCMs for human-auditable environments, and formula-grounded SCMs to test discovery under explicit scientific mechanisms, along with real-world datasets for external validity. Experiments show that performance rankings vary significantly across SCM families and protocols, indicating that strong results on one benchmark do not generalize to others, especially in the context of causal discovery foundation models.

arXiv Machine Learning
Sep 11

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.

By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv AI
Aug 20

CausalProfiler: Generating Synthetic Benchmarks for Rigorous and Transparent Evaluation of Causal Machine Learning

CausalProfiler is a synthetic benchmark generator designed to evaluate causal machine learning (Causal ML) methods more rigorously and transparently. It randomly samples causal models, data, queries, and ground truths based on explicit design choices across observation, intervention, and counterfactual reasoning levels, providing coverage guarantees and transparent assumptions. The authors demonstrate its utility by testing several state‑of‑the‑art methods under diverse conditions, both within and outside the identification regime, highlighting the insights CausalProfiler can reveal.

By Panayiotis Panayiotou, Audrey Poinsot, Alessandro Leite, Nicolas Chesneau, Marc Schoenauer, \"Ozg\"ur \c{S}im\c{s}ek
arXiv Machine Learning
Aug 19

TabCausal: Pretraining Across Causal Environments for Tabular Causal Discovery

TabCausal is a causal discovery foundation model that learns to map datasets directly to causal graphs by pretraining across diverse causal environments. It uses a dynamic task construction strategy to expose the model to varied graph priors, mechanisms, noise models, dimensions, sample sizes, and intervention regimes, improving transferability from observational and mixed‑interventional data. On large synthetic benchmarks and a new protocol‑guided semantic benchmark, TabCausal outperforms many classical baselines and shows robust structure recovery, especially when interventional evidence is available.

By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv Machine Learning
Jun 26

Use What You Know: Causal Foundation Models with Partial Graphs

arXiv:2602. 14972v2 Announce Type: replace Abstract: Estimating causal quantities traditionally relies on bespoke estimators tailored to specific assumptions.

By Arik Reuter, Anish Dhir, Cristiana Diaconu, Jake Robertson, Ole Ossen, Frank Hutter, Adrian Weller, Mark van der Wilk, Bernhard Sch\"olkopf
Hugging Face Trending Papers
Sep 3

Federated Causal Discovery via Regression-Directed Cumulants

The paper investigates federated learning for linear non‑Gaussian acyclic models (LiNGAM), proposing the FedRCD family of algorithms that use higher‑order cumulants to enable privacy‑preserving causal discovery across distributed clients. It addresses limitations of existing federated methods, such as FedISHC’s failure under near‑symmetric noise, and introduces variants that balance communication rounds with algebraic noise handling. Experiments reveal that cumulant‑based federated approaches rank variables by a variance ladder induced by the DAG rather than by population asymmetry, and that marginal standardisation degrades performance while scale‑invariant DirectLiNGAM remains robust.

arXiv Machine Learning
Jul 14

FedCausal-Dyn: A Causal-Dynamic Paradigm for Federated Learning under Dynamic Feature Drift

arXiv:2607. 09695v1 Announce Type: new Abstract: This paper addresses the challenging problem of dynamic feature drift in federated learning, where data distributions evolve across clients and over time -- a common scenario in real-world applications like financial technology.

By Kaijie Chen, Alex Johnson, Maria Garcia, Wei Zhang, Daniel Kim