arXiv AI

Personalized State-Transition-Aware Memory for Clinical Agents

The paper introduces STAM, a state‑transition‑aware memory framework for large language model agents that process clinical records. STAM records changes in a patient’s state as new entries arrive, using semantic retrieval and typed clinical relations to separate current information (Active) from superseded or resolved information (History). During retrieval, a query‑dependent gate selects the appropriate historical memory, enabling accurate question answering and state‑maintenance diagnostics across four longitudinal clinical benchmarks.

arXiv Computation and Language
6d ago

PIA: A Personal Intelligence Agent Turning Health Conversations into Records and Records into Understanding

PIA is a personal intelligence agent that works alongside a consumer health agent to convert health conversations into structured clinical records and to transform those records into a synthesized understanding of the user. It uses a memory system with four controls—extraction, memory, retrieval, and understanding—each supported by a health module that includes a schema, medical alias dictionary, knowledge graph, and temporal rules. The agent demonstrates that deeper memory injection—from simple recall to a health snapshot to a causal trajectory—yields progressively richer answers, while also revealing challenges such as missing self‑reported data, the influence of question phrasing, and the presence of structural noise in causal links.

By Jeonghun Yoon, Dongchan Kim, Hongyeon Yu, Young-Bum Kim, Jaegul Choo
Hugging Face Trending Papers
Sep 24

A Living Benchmark for Information Retrieval from Electronic Health Records

The paper introduces BRIE, a scalable framework that automatically creates question–answer pairs from longitudinal electronic health record notes, validated by nineteen clinicians. It offers a continuously maintainable benchmark for evaluating large language models in clinical settings, addressing limitations of manual, costly, and quickly outdated existing benchmarks. Experiments across nine LLMs and five inference strategies reveal that even state‑of‑the‑art systems often miss clinically important information, especially for synthesis‑heavy queries.

arXiv AI
Sep 25

A Living Benchmark for Information Retrieval from Electronic Health Records

The paper introduces BRIE, a continuously maintainable benchmark for evaluating large language models (LLMs) in electronic health record (EHR) information retrieval. It presents a scalable framework that automatically generates question–answer pairs from longitudinal EHR notes, validated by nineteen clinicians. The benchmark allows assessment of multiple inference strategies and highlights that state‑of‑the‑art LLMs often miss clinically important information, especially when synthesis across documents is required.

By Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani, Philip Chung, Kevin R Keet, Kameron C. Black, Andrea T. Fisher, Sarita Khemani, Jerry Liu, Stephen Ma, Saloni K. Maharaj, Rita M. Pandya, Eduardo Perez-Guerrero, Priyanka Pillai, Lisa Shieh, David J. H. Wu, James Xie, James C. McAvoy, Teresa Nguyen, Jessica Tran, Lucy Yin, Bridget Lin, Alison Callahan, Jason A. Fries, Nigam H. Shah, Emily Alsentzer