arXiv AI By Rahul Gorijavolu, Kaushik Madapati, Pritika Vig, Rawan Abulibdeh, Nikhil Jaiswal, Mahri Kadyrova, Zeamanuel Hailu Tesfaye, Charles Senteio, Paula Maurutto, Leo Anthony Celi

Testing the Black Box: Structural Barriers to Independent Evaluation of Consumer-Facing Health LLMs

Read the original on arXiv AI →

arXiv:2606. 08483v1 Announce Type: new Abstract: Background: Consumer-facing large language models are now a common source of health information, and they interpret and personalize responses rather than retrieve them.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 28

Communication styles and reader preferences of LLM- and human-authored COVID-19 information explanations: a case study

The study compares communication styles of large language models (LLMs) and humans in explaining COVID‑19 misinformation, using a dataset of 1,498 fact‑checking claims and 99 blinded reader evaluations. LLM‑generated explanations scored lower on persuasive strategies, certainty, and alignment with social values, yet over 60% of participants preferred LLM content for clarity, completeness, and persuasiveness. The findings suggest that reader preference may not align with traditional measures of communication quality, highlighting both the promise and limits of LLMs in health communication.

By Jiawei Zhou, Kritika Venkatachalam, Minje Choi, Koustuv Saha, Munmun De Choudhury
arXiv AI
Jul 9

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

arXiv:2601. 16529v4 Announce Type: replace Abstract: Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-based guidelines.

By Dongshen Peng, Yi Wang, Austin Schoeffler, Sun-ha Hong, Brian Suffoletto, David Kim, Carl Preiksaitis, Christian Rose
arXiv AI
Jun 17

Towards Understanding and Measuring COGNITIVE ATROPHY in LLM Behaviour

arXiv:2606. 18129v1 Announce Type: cross Abstract: Recent incidents involving LLMs used for mental-health support reveal a critical evaluation gap: surface-level safety scores do not capture how models behave across realistic, emotionally sensitive interactions over time.

By Abeer Badawi, Moyosoreoluwa Olatosi, Negin Baghbanzadeh, Laleh Seyyed-Kalantari, Frank Rudzicz, R. Shayna Rosenbaum, Sara Pishdadian, Elham Dolatabadi
arXiv Computation and Language
Sep 22

Evaluating Fine-Tuned and Base Language Models in Maternal and Vaccination Healthcare for African Settings

arXiv:2609.22110v1 Announce Type: new Abstract: Background: Large language models (LLMs) can improve healthcare information delivery in low-resource settings but may produce inaccurate or culturally...

By Abdulquddus Ajibade, Oluwaseun Odunsi, Iyinoluwa Animasaun, Chioma Nwakanma-Akanno, Oluwasegun Oguntuase, Oluwafunke Akinbuwa, Abiodun Adereni