arXiv:2608.22062v1 Announce Type: new
Abstract: Longitudinal clinical event-relation verification determines whether a patient record supports a specified relation among two or more clinical events....
By Xingtao Lin, Yubo Feng, Weixin Liu, Hangqi Ren, Junchao Zhou, Caiwan Sun, You Chen
Grounded Adjudication of Variations across Extracted TimeLines (GAVEL) is a new LLM‑based protocol that compares two clinical timelines against their source case report, identifying discrepancy types, issuing verdicts, and citing relevant report passages for each difference. In a study of 126 reports, GAVEL evaluated 2,738 findings from GPT‑5.6sol and DeepSeek V3.2, ranked six LLM extractors and two human annotators, and guided a merging process that improved timeline accuracy—reducing discrepancies from 7.63 to 0.85 per report and yielding a 77.0% preference rate for merged timelines. The approach demonstrates that report‑based comparison can refine extracted timelines without assuming any single timeline as ground truth.
By Jack Cummins, Sayantan Kumar, Ketan Tamirisa, Jeremy C. Weiss
arXiv:2608. 07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably.
By Veronica Chatrath, Bryan Zhu, George Pu, Jingxuan Fan, Apaar Shanker, Varun Ursekar, Anahita Sharma, Jason Qin, Keqi Han, Soham Dinesh Tiwari, Soham Dan, Vijay Kalmath, Yuan Li, Daniel Yue Zhang, Chenguang Wang, Zainab Doctor, Zhijun Yin, Nigam H. Shah, Yuan Xue
MedFabric is a new benchmark for detecting word‑level medical fabrications, comprising 646 fabricated statements each paired with a ground‑truth passage that shares the same LLM authorship and nearly identical wording. The study shows that current detectors perform poorly—expert clinicians achieve only 53.3% macro‑F1 and no detector family surpasses 60% without gold evidence—highlighting that detection hinges on evidence correctness rather than subtlety of fabrication. The authors demonstrate that a retrieval‑confidence gate can substantially improve performance, raising macro‑F1 from 61% to 74%.
By Tung Sum Thomas Kwok, Qian Qian, Xiaofeng Lin, Dongxu Zhang, Jun Han, Zhichao Yang, Davin Hill, Tamer Soliman, Sanjit Singh Batra, Robert Tillman, Guang Cheng
The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.
By Guangzhe Zhang
arXiv:2608. 08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said.
By Fengrong Wan, Chengcan Wu, Ningtao Lyu
arXiv:2609.21387v1 Announce Type: cross
Abstract: Relexicalization is a pivotal technique in clinical NLP, as it facilitates robust masking of sensitive information while synthesizing datasets that r...
By Dipankar Das, Atri Mandal, Sandeep Singh, Tushar Shandhilya
arXiv:2608.31016v1 Announce Type: cross
Abstract: Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the n...
By Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
arXiv:2606. 18037v1 Announce Type: new Abstract: Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, APIs, databases, clinical records, and formulary tools.
By Ander Alvarez, Santhiya Rajan, Samuel Mugel, Rom\'an Or\'us
arXiv:2609.01111v1 Announce Type: new
Abstract: Clinical LLM assistants must reason over multi-visit patient trajectories, yet whether the compact history representations used to scale them---retriev...
By Huimin Wang, Zhengyi Zhao, Yutian Zhao
arXiv:2607. 21859v2 Announce Type: replace Abstract: Constructing causal directed acyclic graphs (DAGs) is a core step in biomedical causal analysis, yet it remains a largely manual process.
By Yi-han Sheu, Michael R. Steigman, Yu Zhou, Bo Wang, Fan-Yu Yen, Jordan W. Smoller
The paper introduces "stale‑document poisoning," a temporal alignment failure where outdated retrieval evidence causes retrieval‑augmented generation models to produce incorrect answers even when the model could answer correctly without retrieval. A benchmark of 317 verified knowledge reversals across medicine, law, software, and platform policy shows that outdated evidence flips 30–37% of Llama and Qwen answers, rising to 66–75% when models are explicitly instructed to trust the document. The study demonstrates that providing explicit validity information and a recency‑aware re‑ranker can substantially reduce poisoning, highlighting the need for models to assess whether retrieved evidence remains applicable.
By Md Shamim Ahmed, Lukas Galke Poech, Richard R\"{o}ttger