arXiv AI

How Clinicians Think and What AI Can Learn From It

arXiv AI
Aug 18

Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning

arXiv:2608. 14804v1 Announce Type: new Abstract: Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient.

By Augusto Bernardo Pissarra, Victor Lorena de Farias Souza
arXiv AI
Sep 2

AI Morbidity and Mortality: A Framework for Clinical AI Failure Review

AI Morbidity and Mortality (AI M&M) is a structured, blameless framework designed to review clinical AI failures. It combines standardized case intake, evidence preservation, investigator reconstruction, tool‑in‑loop attribution, and corrective‑action tracking, classifying each event across four linked dimensions: Trigger, Mechanism, Clinical Pathway, and Corrective Action. The authors demonstrate the framework with five outpatient medication and clinical decision‑support cases, achieving full agreement among reviewers on all classification axes.

By Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu
arXiv AI
Sep 1

INTERVenE: Temporal-Abstraction-Interval Based Transformers for Short-Horizon Medical Event Prediction

INTERVenE introduces Transformer models that use a knowledge‑based temporal abstraction (KBTA) token stream of named clinical concepts instead of raw measurements, enabling per‑token attributions to resolve directly to clinical concepts. Two variants are offered: an auto‑regressive decoder that generates future abstraction trajectories with step‑wise risk readouts, and a bidirectional encoder that jointly predicts risk and time‑to‑event in a single pass. On 57,078 MIMIC‑IV admissions, the encoder variant outperforms neural baselines with a support‑weighted AUPRC of 0.672 and AUROC of 0.901, while the decoder provides complementary token‑level risk trajectories.

By Shahar Oded, Yuval Shahar
arXiv AI
Sep 4

Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications

The paper reviews how Large Language Models (LLMs) are being adapted for medical reasoning, moving beyond single-step answers to systems that can systematically, transparently, and verifiably reason. It introduces a taxonomy of enhancement techniques, split into training-time methods such as supervised fine‑tuning and reinforcement learning, and test-time methods like prompt engineering and multi‑agent systems. The review examines their application across text, image, and code modalities in key clinical areas—diagnosis, education, and treatment planning—and tracks the shift in evaluation benchmarks from simple accuracy to more nuanced assessments of reasoning quality and visual interpretability.

By Zizhan Ma, Wenxuan Wang, Meidan Ding, Shiyi Zheng, Shengyuan Liu, Jie Liu, Jiaming Ji, Linlin Shen, Yixuan Yuan, Wenting Chen