arXiv Machine Learning

CHAIN: Calibrated LLM Forecasting via Causal-Temporal Hypergraph Inference

arXiv Machine Learning
6d ago

BridgeMem: Causal Dyadic Transition Residuals for Temporal Knowledge Graph Forecasting

BridgeMem is a new method for temporal knowledge graph forecasting that focuses on pair‑specific transition evidence, adding a residual correction to the log scores of a frozen full‑vocabulary forecaster. It retrieves and encodes prior events between a query actor and candidate, converting them into a likelihood‑ratio correction via a support‑adaptive empirical‑Bayes reader. Across five benchmarks, BridgeMem outperforms nine baselines from 2021–2026, improving filtered MRR and Hits@{1,3,10} metrics by up to 0.0216.

By Zeyan Li, Libing Chen, Shengda Zhuo, Yin Tang, Jianfeng Xu
arXiv AI
Jun 3

CauTion: Knowing When to Trust LLMs for Ensemble Causal Discovery

arXiv:2606. 03602v1 Announce Type: cross Abstract: Causal discovery from observational data remains challenging due to the fundamental limitations of purely statistical methods, such as statistical distinguishability within equivalence classes and sensitivity to finite sample sizes.

By Bo Peng, Kaiwen Wu, Sirui Chen, Zhiheng Wang, Yu Qiao, Chaochao Lu
arXiv AI
Aug 26

From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

The study evaluates 12 instruction‑tuned open‑weight LLMs on six causal‑graph benchmarks, testing five prompting strategies and four confidence sources. Findings show that LLMs tend to over‑predict edges, misclassify indirect or reversed edges as direct, and exhibit high over‑confidence, while conventional confidence estimates are unreliable and agreement signals offer limited improvement. The results suggest LLMs should be used as externally validated soft causal priors rather than definitive causal‑structure evidence.

By Amit Kumar, Elnur Adl Zarabi, Suranjana Trivedy, Zhiqian Chen, Lei Zhang, Kaiqun Fu, Taoran Ji