arXiv AI

TIDE 2.0: an open, model-agnostic engine for keyed de-identification of clinical notes

Hugging Face Trending Papers
3d ago

Auditable Clinical Timeline Reconstruction with Provenance-Aware Evidence Graphs

The study evaluates an auditable patient‑timeline reconstruction system that tracks provenance, records revisions, and refuses to answer when evidence is missing. Using a synthetic corpus of 1,000 patients and 3,353 notes, two provenance‑aware Evidence Graph operators reduced graph size by 33–37% while preserving all answers across 6,813 query points; a fixed‑window baseline failed to answer over half of the points. The system’s evidence‑gating mechanisms (BioClinicalBERT and a zero‑shot LLM) responded appropriately to evidence‑unavailable controls, but performance varied on marker‑free controls, with BERT maintaining high accuracy but the LLM’s coverage dropping sharply. "whyItMatters":"The results demonstrate that provenance‑aware evidence graphs can significantly reduce data complexity while maintaining answer integrity, highlighting a practical approach to building auditable clinical NLP systems."

arXiv AI
Aug 20

Key Coverage Matters: Semi-Structured Extraction of OCR Clinical Reports

The paper presents a method for extracting key information from OCR‑digitized clinical reports, addressing challenges posed by heterogeneous documents and noisy OCR output. It introduces an open key space that is iteratively mined, normalized, clustered, and verified to build a canonical key inventory, and defines key coverage as a metric for inventory completeness. Experiments on reports from over 20 hospitals using a 0.2B BERT model show that performance improves steadily with key coverage, achieving high F1 scores when the top 90 keys are covered and outperforming a fine‑tuned Qwen3‑0.6B baseline.

By Yu Wang, Yingyun Li, Ying Qin, Haiyang Qian
arXiv Computation and Language
Aug 25

Clinically Grounded Privacy Evaluation of Medical LMs

The paper introduces a clinically grounded privacy evaluation framework for medical language models, assessing leakage across a spectrum of adversarial access levels—from publicly inferable demographics to leaked note fragments. Using this framework on an LM pretrained on 378,000 clinical notes, the authors find that routine encounter metadata leads to high verbatim memorization and significant recovery of sensitive diagnoses (e.g., AUROC 0.91 for abortion, 0.82 for HIV). They also note that exact-match memorization can overstate disclosure, with 36% of memorized tokens being templated documentation, underscoring the risks of training on longitudinal clinical data and offering a reusable evaluation tool.

By Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle, Vivian Utti, Jordan Li Cahoon, Nathaniel Hendrix, Ayin Vala, Marzyeh Ghassemi, Emily Alsentzer
arXiv AI
Sep 1

Review Before Trust: Source-Grounded Integrity Gates for AI-Assisted Personal Health Records

The paper introduces a source‑grounded integrity gate for AI‑assisted personal health records, ensuring that data generated by large language models remains provisional until a deterministic monitor verifies it against the source document. The monitor only accepts candidates that contain a unique supporting quotation, appear within the same laboratory row, and preserve provenance, preventing the model from approving its own output. In Medical DataCloud, the system passed all 22 conformance and mutation tests and, in a replay of nine historical lab reports, admitted 72 of 97 numeric candidates while retaining 25 for human review.

By Nora Girda, Adrian Groza
arXiv Computation and Language
Sep 10

MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

MedDeID is an on‑premises framework that combines in‑house annotation, synthetic‑note generation, model training, inference, pseudonymisation and evaluation to de‑identify clinical text. On a Dutch hospital benchmark, a hospital‑trained transformer detected 98.9 % of identifying text while redacting only 0.24 % of non‑identifier text; a synthetic‑only model achieved 96.1 %. In primary‑care notes, the synthetic‑trained model outperformed the hospital‑trained model in recall and robustness to identifier‑format changes, and an English version trained without real text reached 99.7 % and 98.9 % detection on synthetic benchmarks.

By Stig Hellemans, Tom Stroobants, Elyne Scheurwegs, Pieter Meysman, Philippe G. Jorens, Kris Laukens
arXiv Computation and Language
Sep 11

INDRA: A New AI Tool for Exploring Tobacco, Fossil Fuel, and Chemical Industry Archives

INDRA is a research platform that integrates multiple archival collections—such as UCSF’s Industry Documents Library, Columbia and CUNY’s ToxicDocs, and Stanford’s SRITA—into a single, LLM‑readable corpus. It employs three safeguards: a closed evidentiary sandbox, real‑time provenance tagging, and a deterministic system‑level protocol to ensure that model outputs are clearly distinguished from archival evidence and from the model’s own inferences. The platform enables large‑language‑model‑powered investigations across these archives while keeping the conditions of knowledge production transparent and auditable.

By Daniel Akselrad, Robert N. Proctor
arXiv AI
Aug 19

Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

The study evaluates whether large language models (LLMs) with in‑context learning can better identify institution‑specific protected health information (PHI) in electronic health records than existing de‑identification systems. Using 100 pediatric oncology notes from Texas Children’s Hospital, eight LLMs were compared to two purpose‑built systems and pattern‑based baselines under three progressively specific prompts. The best LLM achieved an F1 score of 0.918, recovering 79% of previously missed PHI categories and reaching a recall of 0.981 after iterative prompt refinement, demonstrating that calibrated single‑pass prompting can close the institutional PHI gap while balancing precision and recall.

By Daniel Palacios, Matthew Brady Neeley, Angel Adetomike Otto, Shalini Dhamodharan, John P. Woodhouse, Chi-fan Lin, Mark Zobeck, Zhandong Liu, Hyun-Hwan Jeong
arXiv AI
Aug 11

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

arXiv:2608. 07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably.

By Veronica Chatrath, Bryan Zhu, George Pu, Jingxuan Fan, Apaar Shanker, Varun Ursekar, Anahita Sharma, Jason Qin, Keqi Han, Soham Dinesh Tiwari, Soham Dan, Vijay Kalmath, Yuan Li, Daniel Yue Zhang, Chenguang Wang, Zainab Doctor, Zhijun Yin, Nigam H. Shah, Yuan Xue
arXiv AI
6d ago

Clinical Note Bloat Reduction for Efficient LLM Use

The paper introduces TRACE, a method that removes duplicated text—known as note bloat—from clinical notes by leveraging EHR metadata and frequency-based de‑duplication. Across 5.3 million notes from diverse patient cohorts, TRACE eliminated 47.3 % of chart text while preserving information extraction and prediction performance, with only 0.3–6.6 % of removed content being author‑generated. The authors project that applying TRACE could yield net savings of $1.00 M to $13.58 M over three years at a large academic center, depending on model pricing schemes.

By Jordan L. Cahoon, Chloe Stanwyck, Asad Aali, Rachel Madding, Sulaiman S. Somani, Emma Sun, Yixing Jiang, Renumathy Dhanasekaran, Emily Alsentzer