LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
arXiv:2607. 09322v1 Announce Type: new Abstract: In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making.
arXiv:2607. 22566v1 Announce Type: new Abstract: MedLoCoMo is a Medical Long-Context Memory benchmark for patient-specific clinical reasoning over multi-admission medical dialogue.
arXiv:2607. 09322v1 Announce Type: new Abstract: In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making.
arXiv:2606. 26105v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong capabilities in short-context reasoning but degrade in performance over long conversational horizons due to context window limitations and inefficient token usage.
arXiv:2610.00562v1 Announce Type: cross Abstract: Longitudinal clinical reasoning requires large language models (LLMs) to identify and integrate relevant evidence distributed across extended patient...
GLoC-EHR is a multimodal language model that processes electronic health records by combining a fixed-size global memory of the entire patient trajectory with a local memory of selected events. It generates hospital-course summaries and masked concept descriptions, then is fine‑tuned to cite evidence before answering clinical questions, using group relative policy optimization to reward correct, evidence‑supported responses. On MIMIC‑IV outcome tasks, GLoC‑EHR achieves the highest macro AUROC among compared models when answering directly, and maintains strong performance with evidence‑cited reasoning while adding distinct supported findings from the local memory.
arXiv:2508. 14817v2 Announce Type: replace-cross Abstract: Objective: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs).
The paper introduces STAM, a state‑transition‑aware memory framework for large language model agents that process clinical records. STAM records changes in a patient’s state as new entries arrive, using semantic retrieval and typed clinical relations to separate current information (Active) from superseded or resolved information (History). During retrieval, a query‑dependent gate selects the appropriate historical memory, enabling accurate question answering and state‑maintenance diagnostics across four longitudinal clinical benchmarks.
arXiv:2510.03536v3 Announce Type: replace-cross Abstract: Multi-turn medical question answering (QA) aims to model realistic clinical diagnosis, where a doctor gathers patient information across mult...
arXiv:2603. 03292v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit high reasoning capacity in medical question-answering, but their tendency to produce hallucinations and outdated knowledge poses critical risks in healthcare fields.
arXiv:2609.23465v1 Announce Type: new Abstract: Long-horizon conversational memory is especially challenging in multi-actor settings, where relevant evidence is distributed across participants and co...
arXiv:2606. 29503v1 Announce Type: cross Abstract: The verbose context problem occurs when structured concepts have token-inefficient textual representations.
arXiv:2606. 15735v1 Announce Type: cross Abstract: Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making.
arXiv:2508. 01401v2 Announce Type: replace-cross Abstract: Physicians spend significant time documenting clinical encounters, a burden that contributes to professional burnout.