arXiv:2608. 13695v1 Announce Type: cross Abstract: Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users.
By Chialuka Prisca-Mary Onuoha, Bright Etornam Sunu, Rashidat Sikiru
arXiv:2508. 00923v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to answer health-related questions and support healthcare workflows, yet evidence for their safety still relies heavily on static benchmarks that can rapidly become obsolete or be optimized against.
By Jiazhen Pan (Cherise), Bailiang Jian (Cherise), Paul Hager (Cherise), Yundi Zhang (Cherise), Che Liu (Cherise), Friederike Jungmann (Cherise), Hongwei Bran Li (Cherise), Julian Canisius (Cherise), Chenyu You (Cherise), Junde Wu (Cherise), Jiayuan Zhu (Cherise), Fenglin Liu (Cherise), Yuyuan Liu (Cherise), Niklas Bubeck (Cherise), Moritz Knolle (Cherise), Chen (Cherise), Chen (Cherise), Christian Wachinger, Zhenyu Gong, Cheng Ouyang, Georgios Kaissis, Benedikt Wiestler, Daniel Rueckert
HerHealthEval is a controlled evaluation framework that tests multilingual understanding of women's-health communication across English, French, and Modern Standard Arabic. It provides six communicative forms—canonical, clinical, layperson, indirect or hedged, emotionally concerned, and deliberately under-specified—each conveying the same clinical concern except the under-specified form omits details to assess clarification needs. The study evaluates multilingual instruction models and QLoRA-adapted variants on tasks such as concern classification, risk calibration, clarification behavior, parse compliance, and cross-form consistency, revealing that high aggregate accuracy can mask safety-relevant failures and that language-invariant risk labels improve performance.
By Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi, Hana Essam Sayed Ahmed Amrya, Mariam Mousa
arXiv:2609.00319v1 Announce Type: cross
Abstract: Online health information seeking is shifting from keyword search, where users consider a ranked list of links, to conversational systems that compos...
By Phuong Anh Nguyen, Jill Noorily, Matthew Flathers, Haruka Notsu, Laura Ospina-Pinillos, Tommy Nguyen, Samantha Clark, Aoife Keane, Grace Thompson, John Torous
arXiv:2605. 03301v2 Announce Type: replace-cross Abstract: De-identification of clinical text is a prerequisite for the secondary use of electronic health records.
By Jose D. Posada, David Love, Somalee Datta, Priya Desai
arXiv:2606. 19640v1 Announce Type: cross Abstract: AI and large language models (LLMs) have emerged as promising tools to address global mental health challenges.
By Yunkai Xu, Saeed Abdullah