Framework for Grounding Healthcare LLMs in a Causal Knowledge Graph: A Cardiovascular Example Pilot
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608. 15382v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty.
The paper introduces BAR, a Budget‑Aware LLM Reasoning framework that enhances post‑discharge risk prediction by integrating external medical knowledge graphs (KGs). BAR refines KGs into disease‑specific evidence graphs with support scores and provenance, then uses an LLM to plan, navigate, and verify evidence within a patient‑specific budget. Experiments on MIMIC‑III and MIMIC‑IV across eight diseases show BAR improves AUPRC by 3.4 points, raises citation precision from 59.8% to 77.9%, and uses only 62‑65% of the allotted budget.
arXiv:2606. 29876v1 Announce Type: cross Abstract: Modern large language models (LLMs) reach 60-70% diagnostic accuracy on complex clinical case benchmarks, but accuracy alone cannot distinguish stable clinically-grounded reasoning from pattern matching.
arXiv:2608. 02877v1 Announce Type: new Abstract: Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fairness audits of predictive models.
arXiv:2606.22419v3 Announce Type: replace Abstract: A recent Nature Medicine study reports that general-purpose frontier LLMs outperform specialized retrieval-augmented clinical tools on medical benc...
arXiv:2609.37788v1 Announce Type: cross Abstract: Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, draw...