arXiv:2607. 09322v1 Announce Type: new Abstract: In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making.
By Yanzhen Chen, Zihan Xu, Xiaocheng Zhang, Zhiting Fan, Weiqi Zhai, Hongxia Xu, Zuozhu Liu
arXiv:2606. 26105v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong capabilities in short-context reasoning but degrade in performance over long conversational horizons due to context window limitations and inefficient token usage.
By Derek Thomas
arXiv:2610.00562v1 Announce Type: cross
Abstract: Longitudinal clinical reasoning requires large language models (LLMs) to identify and integrate relevant evidence distributed across extended patient...
By Taye Akinrele, Noorbakhsh Amiri Golilarz, Subash Neupane, Sudip Mittal, Shahram Rahimi
GLoC-EHR is a multimodal language model that processes electronic health records by combining a fixed-size global memory of the entire patient trajectory with a local memory of selected events. It generates hospital-course summaries and masked concept descriptions, then is fine‑tuned to cite evidence before answering clinical questions, using group relative policy optimization to reward correct, evidence‑supported responses. On MIMIC‑IV outcome tasks, GLoC‑EHR achieves the highest macro AUROC among compared models when answering directly, and maintains strong performance with evidence‑cited reasoning while adding distinct supported findings from the local memory.
By Chaiho Shin, Kwangsoo Kim
arXiv:2508. 14817v2 Announce Type: replace-cross Abstract: Objective: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs).
By Skatje Myers, Dmitriy Dligach, Timothy A. Miller, Samantha Barr, James Landefeld, Yanjun Gao, Matthew Churpek, Anoop Mayampurath, Majid Afshar
The paper introduces STAM, a state‑transition‑aware memory framework for large language model agents that process clinical records. STAM records changes in a patient’s state as new entries arrive, using semantic retrieval and typed clinical relations to separate current information (Active) from superseded or resolved information (History). During retrieval, a query‑dependent gate selects the appropriate historical memory, enabling accurate question answering and state‑maintenance diagnostics across four longitudinal clinical benchmarks.
By Maryam Haghifam, Zahra Rajabi, Yizhou Sun, Carlos Morato