arXiv AI

From Word Sequences to Behavioral Sequences: Adapting Modeling and Evaluation Paradigms for Longitudinal NLP

arXiv:2601. 07988v2 Announce Type: replace-cross Abstract: While NLP typically treats documents as independent and unordered samples, in longitudinal studies, this assumption rarely holds: documents are nested within authors and ordered in time, forming person-indexed, time-ordered $\textit{behavioral sequences}$.

arXiv Machine Learning
Jun 24

PORTER: Language-Grounded Event Representations for Portable Structured EHR Foundation Models

arXiv:2606. 24102v1 Announce Type: cross Abstract: Most electronic health record (EHR) foundation models encode clinical events as discrete event tokens from a fixed vocabulary and therefore cannot directly represent events containing unseen concepts or new combinations of concepts and attributes such as numeric values.

By Lin Lawrence Guo, Adam Paul Yan, Emily Vettese, Lillian Sung
arXiv AI
Jun 2

Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025

arXiv:2606. 02255v1 Announce Type: cross Abstract: Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and how the annotation process was controlled.

By Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli, Yanran Chen, Christian Greisinger, Lotta Kiefer, Christoph Leiter, Subhadeep Roy, Tewodros Achamaleh, Muhammad Arslan Manzoor, Sebastian Pohl, Yufang Hou, Steffen Eger
arXiv AI
Sep 18

CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives

CliniCIRCA is a modular large‑language‑model framework that reconstructs longitudinal mental‑health patient journeys from raw electronic health record narratives. It temporally classifies clinical events in unstructured discharge summaries without explicit timestamps, producing 15,891 tagged events from 52 summaries and correcting 629 errors to create verified gold‑standard timelines. The framework then generates temporally grounded summaries, compressing each source by 1.52×, and scales to produce 1,000 silver‑standard timelines for training, showing that instruction tuning improves event extraction, temporal tagging, and summarization across models.

By Aiwei Ivy Zhang, Nimra Ishfaq, Mohit Chandra, Santiago Alvarez Lesmes, Adam Coscia, Khatiya Chelidze Moon, Xiaohan Ding, Munmun De Choudhury
arXiv AI
Sep 15

Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment

The paper introduces a paired benchmark to detect hindsight bias in clinical language models by comparing model responses to questions posed at a clinically relevant cutoff versus the full timeline. It uses 171 case reports (40 sepsis, 131 GLP‑1/diabetes) with both human‑annotated and LLM‑generated time‑series data, evaluating accuracy, hindsight trap rate, answer instability rate, and hindsight bias rate. Results show that exposing models to the full timeline consistently increases hindsight bias, while truncating the timeline mitigates bias without sacrificing accuracy.

By Misaki Matsuura, Sayantan Kumar, Ojas Kadam, Jeremy C. Weiss
arXiv Machine Learning
Aug 4

EHR2Path: Comprehensive Pathway-Level Modeling of Longitudinal Patient Trajectories from Multimodal Electronic Health Records

arXiv:2506. 04831v3 Announce Type: replace Abstract: Forecasting how a patient's condition is likely to evolve, including possible deterioration, recovery, treatment needs, and care transitions, could support more proactive and personalized care, but requires modeling heterogeneous and longitudinal electronic health record (EHR) data.

By Chantal Pellegrini, Ege \"Ozsoy, David Bani-Harouni, Matthias Keicher, Nassir Navab