arXiv Computation and Language
Aug 28

DocTalkBN: A Novel Dataset of Expert Telemedicine Conversations in Bengali

DocTalkBN is a large-scale multimodal dataset of authentic expert telemedicine conversations in Bengali, comprising 557.63 hours of paired audio and text, 1,515 multi-turn patient calls, and 10,274 host–doctor question–answer exchanges across 26 medical specialties. The dataset contains 1.7 million tokens and preserves the spontaneity and contextual richness of real medical interactions in a low-resource language. Three downstream tasks—medical triage classification, advice safety evaluation, and medical named entity recognition—are constructed to benchmark large language models and encoder-based baselines, demonstrating DocTalkBN’s practical usefulness for clinically grounded reasoning.

By Anik Saha, Fahmida Sultana Naznin, Sadatul Islam Sadi, Ananya Shahrin Promi, Wahid Al Azad Navid, Rifat Shahriyar
arXiv AI
Jul 9

Health System Scale Semantic Search Across Unstructured Clinical Notes

arXiv:2604. 25605v2 Announce Type: replace-cross Abstract: Introduction: Semantic search, which retrieves documents based on conceptual similarity rather than keywords, offers advantages for retrieval of clinical information.

By Faith Wavinya Mutinda, Spandana Makeneni, Anna Lin, Shivaji Dutta, Irit R. Rasooly, Patrick Dibussolo, Shivani Kamath Belman, Hessam Shahriari, Kevin Murphy, Alex B. Ruan, Barbara H. Chaiyachati, Sanjay Chainani, Robert W. Grundmeier, Scott M. Haag, Jeffrey M. Miller, Heather M. Griffis, Ian M. Campbell
arXiv AI
Jul 9

Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering

arXiv:2607. 06641v1 Announce Type: cross Abstract: Large language models (LLMs) achieve promising results on medical question answering benchmarks, yet their use in public health is constrained by hallucinations and the rapid evolution of official guidance.

By Felix Feldman, Joshua Harris, Timothy Laurence, Leo Loman, Ollie Higgins, Fan Grayson, Poonam Soma, Bethany Pace-Bonello, Michael Borowitz, Toby Nonnenmacher
arXiv Machine Learning
Jun 15

Automatic identification of diagnosis from hospital discharge letters via weakly supervised Natural Language Processing

arXiv:2410. 15051v3 Announce Type: replace-cross Abstract: Identifying patient diagnoses from hospital discharge letters is essential for large-scale cohort selection and epidemiological research, but traditional supervised approaches require extensive manual annotation, which is often impractical for large textual datasets.

By Vittorio Torri, Elisa Barbieri, Anna Cantarutti, Carlo Giaquinto, Francesca Ieva