arXiv AI

Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction

arXiv:2605. 04221v2 Announce Type: replace-cross Abstract: Clinical named entity recognition from dental progress notes is challenging because documentation is highly unstructured, domain-specific, and often privacy-sensitive.

arXiv Machine Learning
Sep 25

Language Specificity vs. Domain Diversity: Benchmarking Transformers for Bangla Medical NER

This study benchmarks transformer models for Bangla medical named entity recognition (NER), comparing BanglaBERT, multilingual BERT (mBERT), XLM‑RoBERTa, and GPT‑4o mini under zero‑shot and few‑shot prompting. Across a full test set of 3,179 samples, fine‑tuned XLM‑RoBERTa achieves a new state‑of‑the‑art F1‑score of 0.5959, while BanglaBERT lags with 0.4937, suggesting that domain diversity outweighs language specificity. The analysis shows high performance on Medicine and Specialist entities (F1 > 0.83) but lower accuracy on Symptoms (F1 0.4367), and demonstrates that fine‑tuned transformers outperform prompt‑only approaches by a factor of 3.76.

By Rakib Abdullah, Md. Maruful Islam Maruf
arXiv Computation and Language
Sep 1

Privacy-Preserving Generation of Clinical Narratives from Medical Terminologies

The paper introduces Term2Note, a method for generating full-length clinical notes under differential privacy constraints. It separates content and form, conditioning note sections on medical terms and applying distinct DP protections to terms and notes, followed by a DP quality maximizer. Experiments show that the synthetic notes closely match real clinical notes in statistical properties, and models trained on them perform comparably to those trained on real data, outperforming existing DP text generation baselines.

By Yuping Wu, Viktor Schlegel, Warren Del-Pinto, Srinivasan Nandakumar, Iqra Zahid, Yidan Sun, Hai Li, Usama Farghaly Omar, Amirah Jasmine, Arun-Kumar Kaliya-Perumal, Chun Shen Tham, Gabriel Connors, Anil A Bharath, Goran Nenadic
arXiv Computation and Language
Sep 10

MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

MedDeID is an on‑premises framework that combines in‑house annotation, synthetic‑note generation, model training, inference, pseudonymisation and evaluation to de‑identify clinical text. On a Dutch hospital benchmark, a hospital‑trained transformer detected 98.9 % of identifying text while redacting only 0.24 % of non‑identifier text; a synthetic‑only model achieved 96.1 %. In primary‑care notes, the synthetic‑trained model outperformed the hospital‑trained model in recall and robustness to identifier‑format changes, and an English version trained without real text reached 99.7 % and 98.9 % detection on synthetic benchmarks.

By Stig Hellemans, Tom Stroobants, Elyne Scheurwegs, Pieter Meysman, Philippe G. Jorens, Kris Laukens
arXiv AI
Aug 20

Key Coverage Matters: Semi-Structured Extraction of OCR Clinical Reports

The paper presents a method for extracting key information from OCR‑digitized clinical reports, addressing challenges posed by heterogeneous documents and noisy OCR output. It introduces an open key space that is iteratively mined, normalized, clustered, and verified to build a canonical key inventory, and defines key coverage as a metric for inventory completeness. Experiments on reports from over 20 hospitals using a 0.2B BERT model show that performance improves steadily with key coverage, achieving high F1 scores when the top 90 keys are covered and outperforming a fine‑tuned Qwen3‑0.6B baseline.

By Yu Wang, Yingyun Li, Ying Qin, Haiyang Qian
Hugging Face Trending Papers
Sep 24

A Living Benchmark for Information Retrieval from Electronic Health Records

The paper introduces BRIE, a scalable framework that automatically creates question–answer pairs from longitudinal electronic health record notes, validated by nineteen clinicians. It offers a continuously maintainable benchmark for evaluating large language models in clinical settings, addressing limitations of manual, costly, and quickly outdated existing benchmarks. Experiments across nine LLMs and five inference strategies reveal that even state‑of‑the‑art systems often miss clinically important information, especially for synthesis‑heavy queries.

arXiv AI
Aug 20

MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports

MedStruct‑S is a benchmark for semi‑structured information extraction from OCR‑derived clinical reports, covering key discovery, key‑conditioned QA, and end‑to‑end key‑value extraction. It contains 3,582 annotated real‑world report pages and evaluates models under unknown keys and OCR noise. Experiments show encoder‑only models excel at non‑null key‑conditioned QA, while fine‑tuned decoder‑only models achieve the strongest overall performance across model sizes.

By Yingyun Li, Yu Wang, Haiyang Qian