arXiv:2606. 00018v1 Announce Type: cross Abstract: Ambient AI documentation systems generate clinical note drafts that clinicians frequently revise before signing off into electronic health records, yet how these edits alter hedging language remains unclear.
By Yiliang Zhou, Yawen Guo, Di Hu, Sairam Sutari, Emilie Chow, Steven Tam, Danielle Perret, Deepti Pandita, Kai Zheng
arXiv:2609.22239v1 Announce Type: new
Abstract: Ambient AI is increasingly adopted in healthcare to automatically generate clinical notes from patient-clinician conversations, with the potential to s...
By Jakir Hossain, Yi-Fei Zhao, Hongjian Wang, Minmei Shih, Katie Leigh Mullen, Ahmad P. Tafti, Leming Zhou, Manoj Purohit, William Hogan, Jay Zeng, Elizabeth Skidmore, Yanshan Wang
The paper introduces a benchmark of over 6,000 clinical triage scenarios, 7,000 physician annotations, and 225,000 large language model (LLM) responses to assess how LLMs perform under realistic variations in clinical text. The study finds that LLMs tend to recommend unnecessary care more often than physicians, especially when the input text is perturbed, and that LLM recommendations are more sensitive to gender and tone changes than human recommendations. These findings underscore the importance of deployment‑oriented evaluations that reflect expert physician behavior.
By Abinitha Gourabathina, Haoran Zhang, Yuexing Hao, Walter Gerych, Marzyeh Ghassemi
The paper introduces TRACE, a method that removes duplicated text—known as note bloat—from clinical notes by leveraging EHR metadata and frequency-based de‑duplication. Across 5.3 million notes from diverse patient cohorts, TRACE eliminated 47.3 % of chart text while preserving information extraction and prediction performance, with only 0.3–6.6 % of removed content being author‑generated. The authors project that applying TRACE could yield net savings of $1.00 M to $13.58 M over three years at a large academic center, depending on model pricing schemes.
By Jordan L. Cahoon, Chloe Stanwyck, Asad Aali, Rachel Madding, Sulaiman S. Somani, Emma Sun, Yixing Jiang, Renumathy Dhanasekaran, Emily Alsentzer
MedDeID is an on‑premises framework that combines in‑house annotation, synthetic‑note generation, model training, inference, pseudonymisation and evaluation to de‑identify clinical text. On a Dutch hospital benchmark, a hospital‑trained transformer detected 98.9 % of identifying text while redacting only 0.24 % of non‑identifier text; a synthetic‑only model achieved 96.1 %. In primary‑care notes, the synthetic‑trained model outperformed the hospital‑trained model in recall and robustness to identifier‑format changes, and an English version trained without real text reached 99.7 % and 98.9 % detection on synthetic benchmarks.
By Stig Hellemans, Tom Stroobants, Elyne Scheurwegs, Pieter Meysman, Philippe G. Jorens, Kris Laukens
arXiv:2604. 05435v2 Announce Type: replace Abstract: Incomplete or inconsistent discharge documentation drives care fragmentation and avoidable readmissions.
By Akshat Dasula, Prasanna Desikan, Jaideep Srivastava, Shivali Dalmia, Abhishek Mukherji
arXiv:2606. 05970v1 Announce Type: cross Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream configuration choices is less understood than their accuracy on fixed benchmarks.
By Martin Murin
arXiv:2609.22164v1 Announce Type: cross
Abstract: Structured EHR is abundant but sparse, coded, and difficult to use directly for note-centric clinical modeling. We present MedNotes, a multi-agent sy...
By Nina Fatehi, Reihaneh Hassanzadeh, Meysam Ghaffari, Animesh Agarwal, Carlos Morato
arXiv:2608. 10715v1 Announce Type: cross Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing.
By Lena Holzwarth, Rita Gonz\'alez-M\'arquez, Dmitry Kobak
The paper introduces a clinically grounded privacy evaluation framework for medical language models, assessing leakage across a spectrum of adversarial access levels—from publicly inferable demographics to leaked note fragments. Using this framework on an LM pretrained on 378,000 clinical notes, the authors find that routine encounter metadata leads to high verbatim memorization and significant recovery of sensitive diagnoses (e.g., AUROC 0.91 for abortion, 0.82 for HIV). They also note that exact-match memorization can overstate disclosure, with 36% of memorized tokens being templated documentation, underscoring the risks of training on longitudinal clinical data and offering a reusable evaluation tool.
By Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle, Vivian Utti, Jordan Li Cahoon, Nathaniel Hendrix, Ayin Vala, Marzyeh Ghassemi, Emily Alsentzer
The study examines how professional English editing influences AI text detectors’ false-positive rates for non-native academic writing. Using 135,389 pairs of original and edited manuscripts, researchers found that detector responses varied widely—some editors increased AI scores while others decreased them—and that score changes correlated with the extent of editing. These results highlight professional editing style as a key confounding factor in AI detection, complicating the distinction between AI authorship and linguistic style.
By Hyeonchu Park, Gahye Jeong, Bugeun Kim
arXiv:2606. 26879v1 Announce Type: new Abstract: Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is restricted.
By William Poulett