arXiv AI By Yudai Nakada, Yuichiro Nishiura, Jin Michael Splichal

MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation

Read the original on arXiv AI →

MaSCoD is a multi‑agent framework that explicitly organizes candidate third variables and local structural patterns before making direct‑edge judgments in causal discovery. Using GPT‑5.4 and GPT‑4o on datasets such as Auto‑MPG, DWD, and Sachs, the study shows that providing structural hypotheses in a separate phase (Full) generally improves recall and F1 scores compared to constructing them during judgment, though it can also raise false‑positive rates. Ablation experiments reveal that the benefit of supplying both structural components is not guaranteed, and the advantage varies across datasets and model backbones.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 17

MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation

MaSCoD is a multi‑agent framework that explicitly addresses the premature omission of causal relations by first organizing candidate third variables and local structural patterns before making direct‑edge judgments. Using GPT‑5.4 and GPT‑4o as backbones, the study evaluates MaSCoD on Auto‑MPG, DWD, and Sachs datasets, finding that the Full phase—where structural hypotheses are supplied early—generally improves recall and F1 scores compared to a No Phase 1 approach, though it can also raise false‑positive rates. Ablation experiments reveal that providing both structural and contextual information does not always outperform providing only one, and that the benefits of the Full phase vary across datasets and backbones. whyItMatters":"The paper demonstrates that explicitly pre‑organizing structural context can improve causal graph generation performance, highlighting the importance of designing omission control mechanisms in causal discovery with large language models."

arXiv AI
Aug 26

From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

The study evaluates 12 instruction‑tuned open‑weight LLMs on six causal‑graph benchmarks, testing five prompting strategies and four confidence sources. Findings show that LLMs tend to over‑predict edges, misclassify indirect or reversed edges as direct, and exhibit high over‑confidence, while conventional confidence estimates are unreliable and agreement signals offer limited improvement. The results suggest LLMs should be used as externally validated soft causal priors rather than definitive causal‑structure evidence.

By Amit Kumar, Elnur Adl Zarabi, Suranjana Trivedy, Zhiqian Chen, Lei Zhang, Kaiqun Fu, Taoran Ji
arXiv Machine Learning
Sep 11

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.

By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv Machine Learning
Aug 5

GoT-CD: Graph-of-Thoughts Causal Discovery and the Fragility of Post-hoc Path-Specific Fairness Audits

arXiv:2608. 02877v1 Announce Type: new Abstract: Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fairness audits of predictive models.

By Nitish Nagesh, Elahe Khatibi, Thomas Dean Hughes, Mahdi Bagheri, Pratik Gajane, Amir M. Rahmani
Hugging Face Trending Papers
Sep 10

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena is a unified, evolvable benchmark designed to evaluate causal discovery methods across diverse structural causal models (SCMs). It incorporates synthetic SCMs for controlled structural variation, semantic operational SCMs for human-auditable environments, and formula-grounded SCMs to test discovery under explicit scientific mechanisms, along with real-world datasets for external validity. Experiments show that performance rankings vary significantly across SCM families and protocols, indicating that strong results on one benchmark do not generalize to others, especially in the context of causal discovery foundation models.