arXiv AI By Neil K. R. Sehgal, Dunigan Folk, Lyle Ungar, Sharath Chandra Guntuku

Depression Symptoms and Relational Patterns in 187k ChatGPT Histories

Read the original on arXiv AI →

arXiv:2607. 05685v1 Announce Type: cross Abstract: Large language models are increasingly used as private, always-available conversational systems, but little is known about how people with depressive symptoms use them.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT

arXiv:2609.35953v1 Announce Type: new Abstract: Young people increasingly turn to General-Purpose Conversational Agents (GPCAs), such as ChatGPT, in moments of distress. We examine young adults' (age...

By Marx Wang, Ella Zhang, Cameron Tan, Andrea Mock, Songling Ngo, Zijing Wang, Robert Wolfe, Shirin Amouei, Rachel A. Hanebutt, Desmond C. Ong, Caroline Figueroa, Katie Davis, Anind K. Dey, Alexis Hiniker
arXiv AI
Sep 21

Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation

The paper introduces the COmmunity-centered Peer Engaged Support (COPES) dataset and a three‑axis evaluation framework to gauge how well Large Language Models (LLMs) align with community perspectives on mental‑health support queries. Experiments show that fine‑tuning LLMs on COPES improves strategy alignment and emotion‑tone alignment by over 50% for general‑purpose models, yet these gains are uneven across subreddits and coping strategies. The study also finds that post‑training shifts the model’s recommendations toward problem‑focused advice while reducing emotion‑focused responses, indicating persistent disparities in performance across different communities and needs.

By Mohit Chandra, Nabin Kim, Eli Min, Aamogh Sawant, Tanmay Sutar, Munmun De Choudhury
arXiv Computation and Language
Sep 11

"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated

The study examines how large language models (LLMs) predict depression scores from language responses. In a "Mirror" setup, participants answered structured diagnostic interviews that the LLMs used to predict scores, yielding near-perfect predictions. In a "Non-Mirror" setup, participants gave life history interviews; the LLMs still achieved outstanding prediction accuracy, and both conditions correlated similarly with PHQ-9 scores, indicating that the Mirror advantage disappears when predicting an independent measure. Topic modeling showed different depression themes across interview types, suggesting Mirror evaluations are more about reliability than validity and that Non-Mirror approaches may enhance clinical relevance.

By Tong Li, Rasiq Hussain, Mehak Gupta, Joshua R. Oltmanns