The study evaluates 12 instruction‑tuned open‑weight LLMs on six causal‑graph benchmarks, testing five prompting strategies and four confidence sources. Findings show that LLMs tend to over‑predict edges, misclassify indirect or reversed edges as direct, and exhibit high over‑confidence, while conventional confidence estimates are unreliable and agreement signals offer limited improvement. The results suggest LLMs should be used as externally validated soft causal priors rather than definitive causal‑structure evidence.
By Amit Kumar, Elnur Adl Zarabi, Suranjana Trivedy, Zhiqian Chen, Lei Zhang, Kaiqun Fu, Taoran Ji
The paper investigates whether the factors highlighted by large language models (LLMs) as most influential on their decisions truly reflect necessity or sufficiency in influencing outcomes. By applying controlled black‑box interventions across eight models from Claude, GPT, and Gemini, the authors quantify necessity and sufficiency scores for each factor and compare them to the models’ self‑reported top three factors. Results show modest correlations (≈0.35–0.58) and reveal that the cited top factors often fail to capture the strongest measured influences, indicating limitations in current explanation practices.
By Urja Pawar, Rajitha Ramanayake, Nabeel Kemal, Ashwin Kandath, Owen O'Neill, Guillaume Bourgeon, Houssem Chatbri
MaSCoD is a multi‑agent framework that explicitly organizes candidate third variables and local structural patterns before making direct‑edge judgments in causal discovery. Using GPT‑5.4 and GPT‑4o on datasets such as Auto‑MPG, DWD, and Sachs, the study shows that providing structural hypotheses in a separate phase (Full) generally improves recall and F1 scores compared to constructing them during judgment, though it can also raise false‑positive rates. Ablation experiments reveal that the benefit of supplying both structural components is not guaranteed, and the advantage varies across datasets and model backbones.
By Yudai Nakada, Yuichiro Nishiura, Jin Michael Splichal
MaSCoD is a multi‑agent framework that explicitly addresses the premature omission of causal relations by first organizing candidate third variables and local structural patterns before making direct‑edge judgments. Using GPT‑5.4 and GPT‑4o as backbones, the study evaluates MaSCoD on Auto‑MPG, DWD, and Sachs datasets, finding that the Full phase—where structural hypotheses are supplied early—generally improves recall and F1 scores compared to a No Phase 1 approach, though it can also raise false‑positive rates. Ablation experiments reveal that providing both structural and contextual information does not always outperform providing only one, and that the benefits of the Full phase vary across datasets and backbones.
whyItMatters":"The paper demonstrates that explicitly pre‑organizing structural context can improve causal graph generation performance, highlighting the importance of designing omission control mechanisms in causal discovery with large language models."
arXiv:2606. 05403v1 Announce Type: new Abstract: Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions.
By Rohan N. Pradhan, Steve Goley
FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.
By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang