Can LLMs Reason Over Long Horizons? An Empirical Evaluation of Context Strategies for Longitudinal Clinical Reasoning
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 22566v1 Announce Type: new Abstract: MedLoCoMo is a Medical Long-Context Memory benchmark for patient-specific clinical reasoning over multi-admission medical dialogue.
arXiv:2505. 02722v2 Announce Type: replace Abstract: Although large language models (LLMs) have demonstrated impressive reasoning capabilities across general domains, their effectiveness in real-world clinical practice remains limited.
arXiv:2603. 03292v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit high reasoning capacity in medical question-answering, but their tendency to produce hallucinations and outdated knowledge poses critical risks in healthcare fields.
arXiv:2601.03471v4 Announce Type: replace-cross Abstract: Reliable epidemiological reasoning requires synthesizing study evidence to infer disease burden, transmission dynamics, and intervention effe...
arXiv:2610.01938v1 Announce Type: cross Abstract: Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integr...
The paper introduces a paired benchmark to detect hindsight bias in clinical language models by comparing model responses to questions posed at a clinically relevant cutoff versus the full timeline. It uses 171 case reports (40 sepsis, 131 GLP‑1/diabetes) with both human‑annotated and LLM‑generated time‑series data, evaluating accuracy, hindsight trap rate, answer instability rate, and hindsight bias rate. Results show that exposing models to the full timeline consistently increases hindsight bias, while truncating the timeline mitigates bias without sacrificing accuracy.