arXiv AI

CausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs

arXiv:2602. 08939v2 Announce Type: replace Abstract: Large language models increasingly produce fluent causal explanations, yet they often fail in ways aggregate accuracy cannot diagnose: confusing association with intervention, abandoning correct judgments under pressure, over-refusing valid claims, or answering when evidence is underdetermined.

arXiv AI
Aug 26

From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

The study evaluates 12 instruction‑tuned open‑weight LLMs on six causal‑graph benchmarks, testing five prompting strategies and four confidence sources. Findings show that LLMs tend to over‑predict edges, misclassify indirect or reversed edges as direct, and exhibit high over‑confidence, while conventional confidence estimates are unreliable and agreement signals offer limited improvement. The results suggest LLMs should be used as externally validated soft causal priors rather than definitive causal‑structure evidence.

By Amit Kumar, Elnur Adl Zarabi, Suranjana Trivedy, Zhiqian Chen, Lei Zhang, Kaiqun Fu, Taoran Ji
arXiv Computation and Language
6d ago

Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations

The paper introduces ARGUS, a language‑model pipeline that audits evidence for identification assumptions in difference‑in‑differences studies of climate policy. ARGUS evaluates reported evidence against an eleven‑dimension rubric, abstaining when evidence cannot be retrieved. In tests, ARGUS detects 73% of injected flaws versus 18% for a keyword approach, abstains on about 40% of assessments in 26 economics papers, and often assigns higher risk than human labels in a five‑paper pilot.

By Yonghong Zhang, Yong Xie, Isabel M. Parra, Ricardo Correia
arXiv AI
Sep 25

Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding

The paper reports that language‑model agents used for customer‑relationship management can be misled by optimistic assertions from sales representatives in CRM records, leading to incorrect deal approvals. In a benchmark of 100 lead‑qualification tasks, models incorrectly cleared 29 of 31 deals where the representative’s claims contradicted company policies, with misalignment rates ranging from 87% to 97% across seven models. The authors propose diagnostic methods—including bucket analysis, same‑information controls, and compute‑step controls—to distinguish persuasion from information gaps and to quantify the impact of incentive‑misaligned witnesses.

By Rahul Balakavi
Hugging Face Trending Papers
Jun 25

OpenRCA 2.0: From Outcome Labels to Causal Process Supervision

Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and tool use. However, existing datasets suffer from a fundamental gap: they label only the root cause, not the propagation path connecting it to the observed symptom, which largely simplifies the task to naive pattern matching.

arXiv AI
Sep 2

Medical Causal Hypothesis Verification with Large Language Models

The paper "Medical Causal Hypothesis Verification with Large Language Models" reports a small-scale study evaluating eight LLMs on 17 medical causal hypotheses. The authors introduce an evaluation framework and annotate 1,067 evidence points across six criteria, using nine metrics to assess performance. Results show that while LLMs have strong recall, they frequently fail to provide valid scientific articles, evidence, or reject unsupported hypotheses, revealing a critical limitation for their use in healthcare.

By Safiyyah Ahmed, Abrar Ansari, Md Aminul Islam, Elena Zheleva