arXiv AI

Understanding Stigmatizing Language in Clinical Documentation: A Paired Comparison of Ambient AI Drafts and Clinician Finalized Notes

arXiv:2606. 00019v1 Announce Type: cross Abstract: Ambient artificial intelligence (AI) documentation tools are increasingly deployed to reduce clinician documentation burden, but their implications for biased language in clinical notes remain unclear.

arXiv AI
Jun 2

Examine Clinicians' Modification of Hedging Language in Ambient AI Documentation: A Comparative Study of AI Drafts and Final Notes

arXiv:2606. 00018v1 Announce Type: cross Abstract: Ambient AI documentation systems generate clinical note drafts that clinicians frequently revise before signing off into electronic health records, yet how these edits alter hedging language remains unclear.

By Yiliang Zhou, Yawen Guo, Di Hu, Sairam Sutari, Emilie Chow, Steven Tam, Danielle Perret, Deepti Pandita, Kai Zheng
arXiv Computation and Language
Sep 22

Knowledge Graph-Augmented Ambient AI for Clinical Note Generation

arXiv:2609.22239v1 Announce Type: new Abstract: Ambient AI is increasingly adopted in healthcare to automatically generate clinical notes from patient-clinician conversations, with the potential to s...

By Jakir Hossain, Yi-Fei Zhao, Hongjian Wang, Minmei Shih, Katie Leigh Mullen, Ahmad P. Tafti, Leming Zhou, Manoj Purohit, William Hogan, Jay Zeng, Elizabeth Skidmore, Yanshan Wang
arXiv AI
3d ago

Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

The paper introduces a benchmark of over 6,000 clinical triage scenarios, 7,000 physician annotations, and 225,000 large language model (LLM) responses to assess how LLMs perform under realistic variations in clinical text. The study finds that LLMs tend to recommend unnecessary care more often than physicians, especially when the input text is perturbed, and that LLM recommendations are more sensitive to gender and tone changes than human recommendations. These findings underscore the importance of deployment‑oriented evaluations that reflect expert physician behavior.

By Abinitha Gourabathina, Haoran Zhang, Yuexing Hao, Walter Gerych, Marzyeh Ghassemi
arXiv AI
2d ago

Clinical Note Bloat Reduction for Efficient LLM Use

The paper introduces TRACE, a method that removes duplicated text—known as note bloat—from clinical notes by leveraging EHR metadata and frequency-based de‑duplication. Across 5.3 million notes from diverse patient cohorts, TRACE eliminated 47.3 % of chart text while preserving information extraction and prediction performance, with only 0.3–6.6 % of removed content being author‑generated. The authors project that applying TRACE could yield net savings of $1.00 M to $13.58 M over three years at a large academic center, depending on model pricing schemes.

By Jordan L. Cahoon, Chloe Stanwyck, Asad Aali, Rachel Madding, Sulaiman S. Somani, Emma Sun, Yixing Jiang, Renumathy Dhanasekaran, Emily Alsentzer
arXiv Computation and Language
Sep 10

MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

MedDeID is an on‑premises framework that combines in‑house annotation, synthetic‑note generation, model training, inference, pseudonymisation and evaluation to de‑identify clinical text. On a Dutch hospital benchmark, a hospital‑trained transformer detected 98.9 % of identifying text while redacting only 0.24 % of non‑identifier text; a synthetic‑only model achieved 96.1 %. In primary‑care notes, the synthetic‑trained model outperformed the hospital‑trained model in recall and robustness to identifier‑format changes, and an English version trained without real text reached 99.7 % and 98.9 % detection on synthetic benchmarks.

By Stig Hellemans, Tom Stroobants, Elyne Scheurwegs, Pieter Meysman, Philippe G. Jorens, Kris Laukens
arXiv Computation and Language
Aug 25

Clinically Grounded Privacy Evaluation of Medical LMs

The paper introduces a clinically grounded privacy evaluation framework for medical language models, assessing leakage across a spectrum of adversarial access levels—from publicly inferable demographics to leaked note fragments. Using this framework on an LM pretrained on 378,000 clinical notes, the authors find that routine encounter metadata leads to high verbatim memorization and significant recovery of sensitive diagnoses (e.g., AUROC 0.91 for abortion, 0.82 for HIV). They also note that exact-match memorization can overstate disclosure, with 36% of memorized tokens being templated documentation, underscoring the risks of training on longitudinal clinical data and offering a reusable evaluation tool.

By Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle, Vivian Utti, Jordan Li Cahoon, Nathaniel Hendrix, Ayin Vala, Marzyeh Ghassemi, Emily Alsentzer
arXiv AI
Aug 28

Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing

The study examines how professional English editing influences AI text detectors’ false-positive rates for non-native academic writing. Using 135,389 pairs of original and edited manuscripts, researchers found that detector responses varied widely—some editors increased AI scores while others decreased them—and that score changes correlated with the extent of editing. These results highlight professional editing style as a key confounding factor in AI detection, complicating the distinction between AI authorship and linguistic style.

By Hyeonchu Park, Gahye Jeong, Bugeun Kim