The paper evaluates the safety of conversational AI therapy bots for Generation Alpha, revealing that while these models understand 76‑82% of youth‑specific vocabulary, they correctly assess clinical risk only 64‑72% of the time, creating a significant vocabulary‑comprehension gap. Six failure patterns—such as sarcasm masking, minimization acceptance, and semantic drift—were identified, with compounded errors leading to a 94% miss rate when three or more patterns co‑occur. The authors estimate 146,880 missed crises annually and recommend mandatory human‑in‑the‑loop systems, quarterly youth‑specific validation, transparent performance disclosure, and regulatory oversight for youth‑facing mental health AI.
By Manisha Mehta, Virendra Mehta
arXiv:2608.23028v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage th...
By Zeyu Feng, Qingyu Wu, Yuzhe Luo, Hua Cheng
arXiv:2608.30585v1 Announce Type: new
Abstract: Large language models are trained to follow instructions while refusing harmful requests. Jailbreaks exploit this balance to elicit content a model wou...
By Md Mokarram Chowdhury, Ernie Chang, Yang Li
arXiv:2603.19574v2 Announce Type: replace-cross
Abstract: Conversational AI systems are increasingly used for personal reflection and emotional disclosure, raising concerns about their effects on vul...
By Soorya Ram Shimgekar, Vipin Gunda, Jiwon Kim, Violeta J. Rodriguez, Hari Sundaram, Koustuv Saha
arXiv:2608.11025v2 Announce Type: replace
Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A...
By Clemens Vetter, David Kacz\'er, Lucie Flek, Florian Mai
Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies.
arXiv:2409.02244v3 Announce Type: replace-cross
Abstract: Large language models (LLMs) are increasingly being used as ad hoc therapists. While prior research has found that LLMs outperform human coun...
By Zainab Iftikhar, Sean Ransom, Amy Xiao, Nicole Nugent, Jeff Huang
arXiv:2607. 20449v1 Announce Type: cross Abstract: LLMs are trained predominantly on human-authored text, yet the structural and narrative conventions embedded in that text are rarely examined as a source of systematic behavioral influence, or as a governance risk in deployed systems.
By Adam Rigby, Raz Saremi, Azadeh Sohrabinejad, Mehdi Rahimi
The paper introduces StratCBT, a new dataset of 9,688 psychological counseling sessions with 256K utterances, each counselor response aligned to one of eight Cognitive Behavioral Therapy (CBT) strategies. It was created by modeling clients’ negative thoughts and generating high‑quality conversations through self‑chat, using realistic sessions for guidance. Experiments show that strategy‑aligned generation improves professional and effective counseling when evaluated with large language model‑simulated clients.
By Zimu Wang, Yiwen Jiang, Xiangyu Zhao, Yaling Shen, Jiahe Liu, Stephanie Fong, Maxmartwell H Cheng, Guilherme C Oliveira, Anh Nguyen, Robert Desimone, Barnaby Nelson, Dominic Dwyer, Zongyuan Ge
Graph2Counsel is a framework that generates synthetic counseling dialogues by leveraging Client Psychological Graphs (CPGs) to encode the relationships among a client’s thoughts, emotions, and behaviors. The system uses a structured prompting pipeline guided by counselor strategies and explores techniques such as Chain‑of‑Thought and Multi‑Agent Feedback to produce 760 realistic sessions from 76 CPGs. Expert evaluation shows the dataset surpasses previous ones in specificity, counselor competence, authenticity, conversational flow, and safety, and fine‑tuning an open‑source model on it improves performance on several counseling benchmarks.
By Aishik Mandal, Hiba Arnaout, Clarissa W. Ong, Juliet Bockhorst, Kate Sheehan, Rachael Moldow, Tanmoy Chakraborty, Iryna Gurevych
arXiv:2604. 19139v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI, a growing and increasingly conspicuous phenomenon has emerged: the proliferation of verbal tics--repetitive, formulaic linguistic patterns that pervade model outputs.
By Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, Ran Wang
The article surveys how large language models (LLMs) are being applied to mental health, outlining a three‑phase evolution: Phase I uses LLMs as passive information tools and pattern recognizers for assessment; Phase II employs them as empathetic conversationalists for stateless, in‑the‑moment interactions; Phase III aims to create longitudinal, personalized companions that act as stateful cognitive agents. It systematically reviews core technologies, agent architectures (Profile, Memory, Reasoning, Planning), and the datasets and benchmarks that support this progression, offering a coherent narrative and roadmap for future research. The survey also provides a curated resource list at https://github.com/Emo-gml/Awesome-Mental-Health-LLMs.
By He Hu, Yucheng Zhou, Qianning Wang, Yingjian Zou, Chiyuan Ma, Juzheng Si, Jianzhuang Liu, Zitong Yu, Laizhong Cui, Fei Ma, Qi Tian