arXiv Machine Learning
5d ago

BridgeMem: Causal Dyadic Transition Residuals for Temporal Knowledge Graph Forecasting

BridgeMem is a new method for temporal knowledge graph forecasting that focuses on pair‑specific transition evidence, adding a residual correction to the log scores of a frozen full‑vocabulary forecaster. It retrieves and encodes prior events between a query actor and candidate, converting them into a likelihood‑ratio correction via a support‑adaptive empirical‑Bayes reader. Across five benchmarks, BridgeMem outperforms nine baselines from 2021–2026, improving filtered MRR and Hits@{1,3,10} metrics by up to 0.0216.

By Zeyan Li, Libing Chen, Shengda Zhuo, Yin Tang, Jianfeng Xu
arXiv AI
Sep 15

Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment

The paper introduces a paired benchmark to detect hindsight bias in clinical language models by comparing model responses to questions posed at a clinically relevant cutoff versus the full timeline. It uses 171 case reports (40 sepsis, 131 GLP‑1/diabetes) with both human‑annotated and LLM‑generated time‑series data, evaluating accuracy, hindsight trap rate, answer instability rate, and hindsight bias rate. Results show that exposing models to the full timeline consistently increases hindsight bias, while truncating the timeline mitigates bias without sacrificing accuracy.

By Misaki Matsuura, Sayantan Kumar, Ojas Kadam, Jeremy C. Weiss
arXiv Machine Learning
Sep 21

Gradient-Stable Attention Heads Signal LLM Correctness

The paper introduces HeadEntropy, a training‑free method that predicts the correctness of large language model (LLM) answers by measuring how stable each attention head’s pattern is to further gradient updates. By linking the trace of the softmax Jacobian to 2‑Renyi entropy, the authors show that attention spread correlates with gradient stability, enabling accurate hallucination detection without reference annotations. Across five instruction‑tuned LLMs and five diverse datasets—including medicine, multi‑hop reasoning, and mathematics—HeadEntropy achieves a 0.736 AUROC, outperforming other training‑free baselines and matching hidden‑state probes while incurring less than 1% of inference cost.

By Sophie Ostmeier, Brian Axelrod, Maya Varma, Asad Aali, Yabin Zhang, Magdalini Paschali, Sanmi Koyejo, Curtis Langlotz, Akshay Chaudhari
arXiv AI
Aug 19

Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting

The paper introduces ICD-Deepresearch, a workflow that combines foundation models for electronic health records (EHR) and language models with medical search and ICD dictionaries to forecast future ICD codes for upcoming clinical encounters. It evaluates candidate code transitions by linking patient evidence, external clinical relations, and exact code semantics within a fixed top‑K budget, using SparseEHR for initial priors, GPT‑5 for complementary forecasts, and a final selection step that validates, deduplicates, and ranks candidates. The method achieves patient‑averaged precision/recall of 24.60/35.09% on MIMIC‑III and 25.14/48.32% on MIMIC‑IV, with physicians rating 51–68% of its retrieved documents as useful, outperforming standalone GPT‑5 web search and Medical Deep Research.

By Junda Wang, Meysam Ghaffari, Akshat Choube, Mohsen Sharifi Renani, Hong Yu, Carlos Morato