arXiv:2606. 03198v1 Announce Type: cross Abstract: Clinical AI evaluation increasingly delegates scoring to large language models (LLMs) acting as AI raters, yet their scoring behavior across evaluation conditions has not been quantitatively characterized.
By Sangwon Baek, Kyu Yeon Hur, Kyunga Kim
The study introduces MedQADE, a German open‑response clinical benchmark with 3,800 question‑answer pairs and physician reference annotations. It evaluates large language models (LLMs) as judges, finding that while some LLMs (e.g., Gemini 3 Flash) achieve physician‑level agreement on correctness, they exhibit self‑bias and low abstention rates. Physicians showed moderate agreement on correctness but limited agreement on difficulty, and their abstention increased with perceived difficulty.
By William Philipp, Finn Fassbender, Daniel Fister, Thorsten Langer, Martje G. Pauly, Rebecca Herzog, Markus A. Hobert, Theresa Paulus, Alexander Baumann, Chi Wang Ip, Lukas L. Goede, Johanna Reimer, Sebastian L\"ons, Ronald B\"ock, Sebastian Fudickar
arXiv:2608. 08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response burden.
By Yifan Wang
arXiv:2609.15855v1 Announce Type: cross
Abstract: % !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk con...
By Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj, Rachel Francois-Walcott, Aruba Hussain, Sarah Ingram, Angela Loulopoulou, Adva Segal, Elena Volkova
The study examines how large language models (LLMs) predict depression scores from language responses. In a "Mirror" setup, participants answered structured diagnostic interviews that the LLMs used to predict scores, yielding near-perfect predictions. In a "Non-Mirror" setup, participants gave life history interviews; the LLMs still achieved outstanding prediction accuracy, and both conditions correlated similarly with PHQ-9 scores, indicating that the Mirror advantage disappears when predicting an independent measure. Topic modeling showed different depression themes across interview types, suggesting Mirror evaluations are more about reliability than validity and that Non-Mirror approaches may enhance clinical relevance.
By Tong Li, Rasiq Hussain, Mehak Gupta, Joshua R. Oltmanns
arXiv:2607. 18828v1 Announce Type: new Abstract: Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks.
By Koyar Afrasyab