arXiv AI

Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports

Grounded Adjudication of Variations across Extracted TimeLines (GAVEL) is a new LLM‑based protocol that compares two clinical timelines against their source case report, identifying discrepancy types, issuing verdicts, and citing relevant report passages for each difference. In a study of 126 reports, GAVEL evaluated 2,738 findings from GPT‑5.6sol and DeepSeek V3.2, ranked six LLM extractors and two human annotators, and guided a merging process that improved timeline accuracy—reducing discrepancies from 7.63 to 0.85 per report and yielding a 77.0% preference rate for merged timelines. The approach demonstrates that report‑based comparison can refine extracted timelines without assuming any single timeline as ground truth.

arXiv AI
Sep 18

CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives

CliniCIRCA is a modular large‑language‑model framework that reconstructs longitudinal mental‑health patient journeys from raw electronic health record narratives. It temporally classifies clinical events in unstructured discharge summaries without explicit timestamps, producing 15,891 tagged events from 52 summaries and correcting 629 errors to create verified gold‑standard timelines. The framework then generates temporally grounded summaries, compressing each source by 1.52×, and scales to produce 1,000 silver‑standard timelines for training, showing that instruction tuning improves event extraction, temporal tagging, and summarization across models.

By Aiwei Ivy Zhang, Nimra Ishfaq, Mohit Chandra, Santiago Alvarez Lesmes, Adam Coscia, Khatiya Chelidze Moon, Xiaohan Ding, Munmun De Choudhury
arXiv AI
Jun 6

PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage

arXiv:2606. 05463v1 Announce Type: new Abstract: Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically performed manually by patient safety experts.

By Keqi Han, Ryan Young, Annabel Strauss, Lindsey Hughes, Katharine M. Nesbitt, Nicole Schueler, Che Ngufor, Carl Yang, Yuan Xue, Zhijun Yin
arXiv AI
Sep 1

From Analytics to Tumor Boards: An Evidence-Linked Multi-Agent Workflow for Oncology Feature Extraction

arXiv:2608.28974v1 Announce Type: new Abstract: Clinically relevant oncology information is distributed across heterogeneous, longitudinal documentation, creating substantial abstraction burden and r...

By Daniel Kang, Michelle Hu, Soorya Ram Shimgekar, Shayan Vassef, Yufan Wang, Anit Kumar Sahu, Munmun De Choudhury, Vedant Das Swain, Christian Poellabauer, Li Yan Khor, Koustuv Saha, Robert Wojciechowski, Elliot Kidd, Piyum Zonooz, Navin Kumar
arXiv Computation and Language
Aug 24

An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study

The study evaluates large language models (LLMs) on unprocessed electronic medical record data for clinical registry abstraction, focusing on the American College of Cardiology National Cardiovascular Data Registry. In a pilot at one academic center, the LLM identified candidate data sources for each registry question, which abstractors used to define question‑specific document sets. In a subsequent validation at a second center, the LLM answered 157 registry questions with an overall mean accuracy of 91.5%, but accuracy dropped from 96% for simple medication or event flag questions to 62% for event timing questions, reflecting increasing ambiguity and required clinical reasoning.

By James Matheson, Betsy Castillo, Andrew Y. Shin, David Scheinker
arXiv Computation and Language
Sep 7

VERGE: Verification-Enhanced Refinement for Grounded Extraction of Early-Onset Colorectal Cancer Symptoms in Clinical Notes

The paper introduces VERGE, a verification-enhanced refinement workflow that extracts six red‑flag symptoms and family‑history risk status for early‑onset colorectal cancer from free‑text clinical notes. VERGE uses retrieval‑augmented generation followed by a bounded verification‑refinement cycle that checks textual grounding and clinical validity, correcting claims until resolved or escalating to human review. In evaluation on 4,033 clinician‑labeled note‑finding pairs, VERGE improved precision from 0.764 to 0.849 and MCC from 0.681 to 0.730 compared to a single‑agent baseline, while requiring human review for only 1.5 % of claims.

By Nikkie Hooman, Monarch Nigam, Amy E. Hughes, Rasmi G. Nair, Mehak Gupta
arXiv Computation and Language
Sep 25

Clinical Intent Extraction: A FHIR-Aligned Representation and the CIRCA Benchmark

The paper introduces Clinical Intent Extraction (CIE), a task that transforms fragmented clinical action annotations into complete structured records called Clinical Intent Representation (CIR). CIR decomposes each action into verb, type, coded target, timing, condition, request‑intent (aligned to HL7 FHIR) and modality, adding dimensions absent in prior datasets. By re‑expressing five heterogeneous corpora into CIR, the authors create CIRCA, a benchmark of 10,011 harmonized intents with human‑validated subsets, crosswalks, and a deterministic FHIR R4 mapper, and demonstrate that existing models perform poorly on the full task, highlighting the need for targeted development.

By Alexander Apartsin, Yehudit Aperstein