arXiv AI

Testing the Black Box: Structural Barriers to Independent Evaluation of Consumer-Facing Health LLMs

arXiv:2606. 08483v1 Announce Type: new Abstract: Background: Consumer-facing large language models are now a common source of health information, and they interpret and personalize responses rather than retrieve them.

arXiv AI
Aug 28

Communication styles and reader preferences of LLM- and human-authored COVID-19 information explanations: a case study

The study compares communication styles of large language models (LLMs) and humans in explaining COVID‑19 misinformation, using a dataset of 1,498 fact‑checking claims and 99 blinded reader evaluations. LLM‑generated explanations scored lower on persuasive strategies, certainty, and alignment with social values, yet over 60% of participants preferred LLM content for clarity, completeness, and persuasiveness. The findings suggest that reader preference may not align with traditional measures of communication quality, highlighting both the promise and limits of LLMs in health communication.

By Jiawei Zhou, Kritika Venkatachalam, Minje Choi, Koustuv Saha, Munmun De Choudhury
arXiv AI
Jul 9

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

arXiv:2601. 16529v4 Announce Type: replace Abstract: Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-based guidelines.

By Dongshen Peng, Yi Wang, Austin Schoeffler, Sun-ha Hong, Brian Suffoletto, David Kim, Carl Preiksaitis, Christian Rose
arXiv AI
Jun 17

Towards Understanding and Measuring COGNITIVE ATROPHY in LLM Behaviour

arXiv:2606. 18129v1 Announce Type: cross Abstract: Recent incidents involving LLMs used for mental-health support reveal a critical evaluation gap: surface-level safety scores do not capture how models behave across realistic, emotionally sensitive interactions over time.

By Abeer Badawi, Moyosoreoluwa Olatosi, Negin Baghbanzadeh, Laleh Seyyed-Kalantari, Frank Rudzicz, R. Shayna Rosenbaum, Sara Pishdadian, Elham Dolatabadi
arXiv Computation and Language
Sep 22

Evaluating Fine-Tuned and Base Language Models in Maternal and Vaccination Healthcare for African Settings

arXiv:2609.22110v1 Announce Type: new Abstract: Background: Large language models (LLMs) can improve healthcare information delivery in low-resource settings but may produce inaccurate or culturally...

By Abdulquddus Ajibade, Oluwaseun Odunsi, Iyinoluwa Animasaun, Chioma Nwakanma-Akanno, Oluwasegun Oguntuase, Oluwafunke Akinbuwa, Abiodun Adereni
arXiv AI
2d ago

Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

The study evaluates counterfactual bias in ten open‑source large language models (LLMs) for pediatric Emergency Severity Index (ESI) prediction. By creating paired clinical vignettes that differ only in demographic or socioeconomic variables, the authors measure shifts in acuity assignment, finding that counterfactual sensitivity varies widely across model families and sizes. A fine‑tuned Qwen2.5‑7B model exhibited the lowest sensitivity, while larger or medical‑domain models sometimes showed greater shifts, highlighting the need for fairness assessment before clinical deployment.

By Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
arXiv Computation and Language
Sep 18

HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication

HerHealthEval is a controlled evaluation framework that tests multilingual understanding of women's-health communication across English, French, and Modern Standard Arabic. It provides six communicative forms—canonical, clinical, layperson, indirect or hedged, emotionally concerned, and deliberately under-specified—each conveying the same clinical concern except the under-specified form omits details to assess clarification needs. The study evaluates multilingual instruction models and QLoRA-adapted variants on tasks such as concern classification, risk calibration, clarification behavior, parse compliance, and cross-form consistency, revealing that high aggregate accuracy can mask safety-relevant failures and that language-invariant risk labels improve performance.

By Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi, Hana Essam Sayed Ahmed Amrya, Mariam Mousa
arXiv AI
Jun 6

Evaluating the Utility of Personal Health Records in Personalized Health AI

arXiv:2605. 18937v2 Announce Type: replace Abstract: Patient-managed Personal Health Records (PHRs) promises to empower patients to better understand their health; but information in the record is complex, potentially hindering insights.

By Rory Sayres, Kejia Chen, Ayush Jain, Matthew Thompson, Jonathan Richina, Xiang Yin, Jimmy Hu, Fan Zhang, Bob Lou, Mike Sanchez, Ines Mezerreg, Meredith Schreier, Hamsa Subramaniam, I-Ching Lee, Yugang Jia, Daniel Mcduff, Yossi Matias, Avinatan Hassidim, Dale Webster, Yun Liu, Jackie Barr, Quang Duong