arXiv AI By Hiroshi Okumura

When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs

Read the original on arXiv AI →

arXiv:2606. 24370v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into decision-support roles in business and policy contexts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

The study evaluates 12 instruction‑tuned open‑weight LLMs on six causal‑graph benchmarks, testing five prompting strategies and four confidence sources. Findings show that LLMs tend to over‑predict edges, misclassify indirect or reversed edges as direct, and exhibit high over‑confidence, while conventional confidence estimates are unreliable and agreement signals offer limited improvement. The results suggest LLMs should be used as externally validated soft causal priors rather than definitive causal‑structure evidence.

By Amit Kumar, Elnur Adl Zarabi, Suranjana Trivedy, Zhiqian Chen, Lei Zhang, Kaiqun Fu, Taoran Ji
arXiv AI
Sep 7

Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence

The paper investigates whether the factors highlighted by large language models (LLMs) as most influential on their decisions truly reflect necessity or sufficiency in influencing outcomes. By applying controlled black‑box interventions across eight models from Claude, GPT, and Gemini, the authors quantify necessity and sufficiency scores for each factor and compare them to the models’ self‑reported top three factors. Results show modest correlations (≈0.35–0.58) and reveal that the cited top factors often fail to capture the strongest measured influences, indicating limitations in current explanation practices.

By Urja Pawar, Rajitha Ramanayake, Nabeel Kemal, Ashwin Kandath, Owen O'Neill, Guillaume Bourgeon, Houssem Chatbri
arXiv AI
Sep 18

MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation

MaSCoD is a multi‑agent framework that explicitly organizes candidate third variables and local structural patterns before making direct‑edge judgments in causal discovery. Using GPT‑5.4 and GPT‑4o on datasets such as Auto‑MPG, DWD, and Sachs, the study shows that providing structural hypotheses in a separate phase (Full) generally improves recall and F1 scores compared to constructing them during judgment, though it can also raise false‑positive rates. Ablation experiments reveal that the benefit of supplying both structural components is not guaranteed, and the advantage varies across datasets and model backbones.

By Yudai Nakada, Yuichiro Nishiura, Jin Michael Splichal
Hugging Face Trending Papers
Sep 17

MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation

MaSCoD is a multi‑agent framework that explicitly addresses the premature omission of causal relations by first organizing candidate third variables and local structural patterns before making direct‑edge judgments. Using GPT‑5.4 and GPT‑4o as backbones, the study evaluates MaSCoD on Auto‑MPG, DWD, and Sachs datasets, finding that the Full phase—where structural hypotheses are supplied early—generally improves recall and F1 scores compared to a No Phase 1 approach, though it can also raise false‑positive rates. Ablation experiments reveal that providing both structural and contextual information does not always outperform providing only one, and that the benefits of the Full phase vary across datasets and backbones. whyItMatters":"The paper demonstrates that explicitly pre‑organizing structural context can improve causal graph generation performance, highlighting the importance of designing omission control mechanisms in causal discovery with large language models."

arXiv Computation and Language
Aug 27

FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.

By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang