arXiv AI

Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT

arXiv AI
Sep 16

From Momentary Emotion Inference to Sustained Emotion Support: Evaluating a Companion Agent in a Longitudinal Study

The study evaluates PAIR, a theory-based emotion‑regulation companion, over 14 days with 19 participants, analyzing 1,093 sessions. Emotion estimates from the agent aligned better with participants’ self‑reported valence and dominance than arousal, and guided conversations led to higher valence and state‑dependent arousal changes. Participants reported feeling understood, and the perceived helpfulness of guided conversations increased over time, highlighting the role of memory updates and cross‑session personalization in sustained emotional support.

By Kexin Quan, Zijian Ding, Jiaye Yong, Qinshi Zhang, Dong Wang, Jessie Chin
arXiv Computation and Language
Sep 21

Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes

This scoping review examined 48 studies on generative AI chatbots designed to deliver motivational interviewing (MI). It found that most systems were text‑based and disembodied, with about half incorporating dynamic adaptation, and that safety reporting was inconsistent. While user perceptions were generally positive and many studies reported MI‑consistent interactions, evidence for sustained behavioral or functional change remains limited.

By Runze Hu, Jingqi Kong, Yang Yang, Yihang Yang, Jingyao Liu, Haizhou Tang, Shanghang Zhang, Zheng Liu
arXiv AI
Aug 24

When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

The paper evaluates the safety of conversational AI therapy bots for Generation Alpha, revealing that while these models understand 76‑82% of youth‑specific vocabulary, they correctly assess clinical risk only 64‑72% of the time, creating a significant vocabulary‑comprehension gap. Six failure patterns—such as sarcasm masking, minimization acceptance, and semantic drift—were identified, with compounded errors leading to a 94% miss rate when three or more patterns co‑occur. The authors estimate 146,880 missed crises annually and recommend mandatory human‑in‑the‑loop systems, quarterly youth‑specific validation, transparent performance disclosure, and regulatory oversight for youth‑facing mental health AI.

By Manisha Mehta, Virendra Mehta
arXiv AI
Sep 15

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

arXiv:2609.15855v1 Announce Type: cross Abstract: % !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk con...

By Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj, Rachel Francois-Walcott, Aruba Hussain, Sarah Ingram, Angela Loulopoulou, Adva Segal, Elena Volkova