MaSCoD is a multi‑agent framework that explicitly organizes candidate third variables and local structural patterns before making direct‑edge judgments in causal discovery. Using GPT‑5.4 and GPT‑4o on datasets such as Auto‑MPG, DWD, and Sachs, the study shows that providing structural hypotheses in a separate phase (Full) generally improves recall and F1 scores compared to constructing them during judgment, though it can also raise false‑positive rates. Ablation experiments reveal that the benefit of supplying both structural components is not guaranteed, and the advantage varies across datasets and model backbones.
By Yudai Nakada, Yuichiro Nishiura, Jin Michael Splichal
The study evaluates 12 instruction‑tuned open‑weight LLMs on six causal‑graph benchmarks, testing five prompting strategies and four confidence sources. Findings show that LLMs tend to over‑predict edges, misclassify indirect or reversed edges as direct, and exhibit high over‑confidence, while conventional confidence estimates are unreliable and agreement signals offer limited improvement. The results suggest LLMs should be used as externally validated soft causal priors rather than definitive causal‑structure evidence.
By Amit Kumar, Elnur Adl Zarabi, Suranjana Trivedy, Zhiqian Chen, Lei Zhang, Kaiqun Fu, Taoran Ji
arXiv:2606. 24370v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into decision-support roles in business and policy contexts.
By Hiroshi Okumura
CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.
By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
CausalArena is a unified, evolvable benchmark designed to evaluate causal discovery methods across diverse structural causal models (SCMs). It incorporates synthetic SCMs for controlled structural variation, semantic operational SCMs for human-auditable environments, and formula-grounded SCMs to test discovery under explicit scientific mechanisms, along with real-world datasets for external validity. Experiments show that performance rankings vary significantly across SCM families and protocols, indicating that strong results on one benchmark do not generalize to others, especially in the context of causal discovery foundation models.
arXiv:2608. 02877v1 Announce Type: new Abstract: Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fairness audits of predictive models.
By Nitish Nagesh, Elahe Khatibi, Thomas Dean Hughes, Mahdi Bagheri, Pratik Gajane, Amir M. Rahmani
arXiv:2404.06349v3 Announce Type: replace
Abstract: The ability to understand causality significantly impacts the competence of large language models (LLMs) in output explanation and counterfactual r...
By Yu Zhou, Xingyu Wu, Jibin Wu, Liang Feng, Kay Chen Tan
arXiv:2609.22409v1 Announce Type: new
Abstract: Understanding contextual causality is critical for large language models (LLMs), as it enables them to accurately identify causal relations in specific...
By Yiheng Zhao, Jun Yan, Chengming Hu
arXiv:2603. 16475v2 Announce Type: replace Abstract: In schema-guided reasoning (SGR) pipelines, LLMs produce explicit intermediate structures -- rubrics, checklists, or verification queries -- before committing to a final decision.
By Oleg Somov, Mikhail Chaichuk, Gleb Ershov, Karim Vafin, Mikhail Seleznyov, Alexander Panchenko, Elena Tutubalina
arXiv:2609.37446v1 Announce Type: new
Abstract: Supervised causal discovery learns to infer causal structure for a new dataset from training datasets paired with structural labels. These training pai...
By Pingchuan Ma, Rui Ding, Bojun Huang, Shuai Wang
arXiv:2608. 03868v1 Announce Type: cross Abstract: Causal Discovery (CD) from observational data faces two fundamental challenges.
By Abhinav Thorat, Ravi Kumar Kolla, Vishak K Bhat, Harsh Vardhan Singh Chauhan, Niranjan Pedanekar
CIDER-FM is a causal foundation model that combines finite observational data with surrogate-interventional datasets to predict target conditional interventional distributions more accurately than using observational data alone. It employs an intervention-aware representation and hierarchical three‑axis attention to integrate information across variables, samples, and experimental regimes. Experiments on synthetic graphs, simulated data, and real‑world Causal Chambers data show that incorporating experimental context improves CID prediction performance.
By Yuche Gao, Arik Reuter, Siyuan Guo, Anish Dhir, Bernhard Sch\"olkopf, Adrian Weller