arXiv AI By Baihan Lin

Language-model ratings of depression reflect the rater more than the patient

Read the original on arXiv AI →

The study examined how language‑model raters assess depression using the Patient Health Questionnaire across 880 raters and 189 interviews. Model choice accounted for 30% of symptom‑score variance, while stable participant differences explained 10.5%. Even raters with similar overall accuracy (AUC ≥ 0.70) disagreed on screening decisions for 40% of participants, and only recalibration with labeled data improved agreement and accuracy modestly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
2d ago

Evaluating Large Language Model Raters for German Open-Response Clinical Questions: A Physician-Annotated Benchmark Study of Agreement, Evaluator Bias, and Abstention

The study introduces MedQADE, a German open‑response clinical benchmark with 3,800 question‑answer pairs and physician reference annotations. It evaluates large language models (LLMs) as judges, finding that while some LLMs (e.g., Gemini 3 Flash) achieve physician‑level agreement on correctness, they exhibit self‑bias and low abstention rates. Physicians showed moderate agreement on correctness but limited agreement on difficulty, and their abstention increased with perceived difficulty.

By William Philipp, Finn Fassbender, Daniel Fister, Thorsten Langer, Martje G. Pauly, Rebecca Herzog, Markus A. Hobert, Theresa Paulus, Alexander Baumann, Chi Wang Ip, Lukas L. Goede, Johanna Reimer, Sebastian L\"ons, Ronald B\"ock, Sebastian Fudickar
arXiv AI
Sep 15

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

arXiv:2609.15855v1 Announce Type: cross Abstract: % !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk con...

By Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj, Rachel Francois-Walcott, Aruba Hussain, Sarah Ingram, Angela Loulopoulou, Adva Segal, Elena Volkova
arXiv Computation and Language
Sep 11

"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated

The study examines how large language models (LLMs) predict depression scores from language responses. In a "Mirror" setup, participants answered structured diagnostic interviews that the LLMs used to predict scores, yielding near-perfect predictions. In a "Non-Mirror" setup, participants gave life history interviews; the LLMs still achieved outstanding prediction accuracy, and both conditions correlated similarly with PHQ-9 scores, indicating that the Mirror advantage disappears when predicting an independent measure. Topic modeling showed different depression themes across interview types, suggesting Mirror evaluations are more about reliability than validity and that Non-Mirror approaches may enhance clinical relevance.

By Tong Li, Rasiq Hussain, Mehak Gupta, Joshua R. Oltmanns