arXiv AI

Consistency evaluation of benchmarks used for causal discovery

arXiv:2606. 01789v1 Announce Type: new Abstract: In graphical causal model, causal discovery aims to construct a causal graph based on numerical data and domain knowledge in plain text.

arXiv Machine Learning
Sep 11

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.

By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
Hugging Face Trending Papers
Sep 10

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena is a unified, evolvable benchmark designed to evaluate causal discovery methods across diverse structural causal models (SCMs). It incorporates synthetic SCMs for controlled structural variation, semantic operational SCMs for human-auditable environments, and formula-grounded SCMs to test discovery under explicit scientific mechanisms, along with real-world datasets for external validity. Experiments show that performance rankings vary significantly across SCM families and protocols, indicating that strong results on one benchmark do not generalize to others, especially in the context of causal discovery foundation models.

arXiv AI
Jun 3

CauTion: Knowing When to Trust LLMs for Ensemble Causal Discovery

arXiv:2606. 03602v1 Announce Type: cross Abstract: Causal discovery from observational data remains challenging due to the fundamental limitations of purely statistical methods, such as statistical distinguishability within equivalence classes and sensitivity to finite sample sizes.

By Bo Peng, Kaiwen Wu, Sirui Chen, Zhiheng Wang, Yu Qiao, Chaochao Lu
arXiv AI
Sep 2

Medical Causal Hypothesis Verification with Large Language Models

The paper "Medical Causal Hypothesis Verification with Large Language Models" reports a small-scale study evaluating eight LLMs on 17 medical causal hypotheses. The authors introduce an evaluation framework and annotate 1,067 evidence points across six criteria, using nine metrics to assess performance. Results show that while LLMs have strong recall, they frequently fail to provide valid scientific articles, evidence, or reject unsupported hypotheses, revealing a critical limitation for their use in healthcare.

By Safiyyah Ahmed, Abrar Ansari, Md Aminul Islam, Elena Zheleva
arXiv Machine Learning
Aug 27

Cluster-Dags as Powerful Background Knowledge For Causal Discovery

The paper introduces Cluster-DAGs as a flexible prior knowledge framework to improve causal discovery. It presents two modified constraint‑based algorithms, Cluster‑PC and Cluster‑FCI, tailored for fully and partially observed data. Experiments on simulated data show that these methods outperform baseline algorithms that lack prior knowledge.

By Jan Marco Ruiz de Vargas, Kirtan Padh, Niki Kilbertus