arXiv Computation and Language

From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health

The article surveys how large language models (LLMs) are being applied to mental health, outlining a three‑phase evolution: Phase I uses LLMs as passive information tools and pattern recognizers for assessment; Phase II employs them as empathetic conversationalists for stateless, in‑the‑moment interactions; Phase III aims to create longitudinal, personalized companions that act as stateful cognitive agents. It systematically reviews core technologies, agent architectures (Profile, Memory, Reasoning, Planning), and the datasets and benchmarks that support this progression, offering a coherent narrative and roadmap for future research. The survey also provides a curated resource list at https://github.com/Emo-gml/Awesome-Mental-Health-LLMs.

arXiv AI
Aug 20

Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges

The paper reviews how large language models are applied in mental health, covering areas such as social media analysis, clinical conversational agents, therapy support tools, prompt engineering, and multimodal learning. It synthesizes interdisciplinary studies that use social media posts, electronic medical records, and multimodal inputs to detect depression, assess suicide risk, provide personalized therapy, and generate psychoeducational content. The review also discusses advances in model interpretability, annotation strategies, multimodal fusion techniques, and highlights ethical, sociotechnical, and regulatory challenges while proposing frameworks for safe, equitable, and accountable deployment.

By Yisong Chen, Yifan Gao, Sijing Yu, Chuqing Zhao, Yang Lu
arXiv AI
Aug 24

When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

The paper evaluates the safety of conversational AI therapy bots for Generation Alpha, revealing that while these models understand 76‑82% of youth‑specific vocabulary, they correctly assess clinical risk only 64‑72% of the time, creating a significant vocabulary‑comprehension gap. Six failure patterns—such as sarcasm masking, minimization acceptance, and semantic drift—were identified, with compounded errors leading to a 94% miss rate when three or more patterns co‑occur. The authors estimate 146,880 missed crises annually and recommend mandatory human‑in‑the‑loop systems, quarterly youth‑specific validation, transparent performance disclosure, and regulatory oversight for youth‑facing mental health AI.

By Manisha Mehta, Virendra Mehta
arXiv AI
2d ago

BiGraph-Diffuse: A Bidirectional Diffusion Language Model with Graph-Structured Retrieval For Mental Health Counseling

BiGraph-Diffuse is a large‑scale diffusion language model designed for mental health counseling, addressing two key limitations of existing AI dialogue systems: the lack of bidirectional understanding for progressive disclosure and the inadequate use of relational clinical knowledge. It pairs this diffusion model with BiGraph‑RAG, a graph‑structured retrieval approach that uses lightweight entity extraction and semantic linking to preserve inferential pathways from symptoms to underlying causes without incurring LLM token costs during indexing. Experiments and theoretical analysis demonstrate the effectiveness of this mutually reinforcing architecture.

By Yuxiang Cheng, Quanwei Tang, Lvhui Lu, Dong Zhang, Shoushan Li, Erik Cambria
arXiv Computation and Language
Aug 31

Roleplaying with Structure: Synthetic Therapist-Client Conversation Generation from Questionnaires

arXiv:2510.25384v2 Announce Type: replace Abstract: Large Language Models (LLMs) are promising tools for synthetic data generation in mental health. However, privacy policies and restrictions forced...

By Doan Nam Long Vu, Rui Tan, Lena Moench, Svenja Jule Francke, Daniel Woiwod, Florian Thomas-Odenthal, Sanna Stroth, Tilo Kircher, Christiane Hermann, Udo Dannlowski, Hamidreza Jamalabadi, Simone Balloccu, Shaoxiong Ji
arXiv Computation and Language
Sep 1

Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs

Graph2Counsel is a framework that generates synthetic counseling dialogues by leveraging Client Psychological Graphs (CPGs) to encode the relationships among a client’s thoughts, emotions, and behaviors. The system uses a structured prompting pipeline guided by counselor strategies and explores techniques such as Chain‑of‑Thought and Multi‑Agent Feedback to produce 760 realistic sessions from 76 CPGs. Expert evaluation shows the dataset surpasses previous ones in specificity, counselor competence, authenticity, conversational flow, and safety, and fine‑tuning an open‑source model on it improves performance on several counseling benchmarks.

By Aishik Mandal, Hiba Arnaout, Clarissa W. Ong, Juliet Bockhorst, Kate Sheehan, Rachael Moldow, Tanmoy Chakraborty, Iryna Gurevych
arXiv Computation and Language
Sep 18

CounselReflect: Opportunities and Challenges for Designing Tools to Support Self-Reflection on Mental Health and Well-Being Conversations with AI

The paper introduces CounselReflect, a tool that converts counseling quality metrics into a framework for users to reflect on their mental‑health AI conversations. Through interviews with 21 users, the study finds that while most participants rarely reflect on their interactions, they identify specific questions they would like such a tool to address. The findings also reveal that users tend to confirm existing beliefs and focus on familiar dimensions, highlighting the need for reflection tools to expose blind spots and encourage a more comprehensive examination of AI interactions, especially when revisiting emotionally charged exchanges.

By Yahan Li, Chaohao Du, Christopher Chun Kuizon, Zeyang Li, Nimra Ishfaq, Shupeng Cheng, Angelica Yinling Sun, Adam C. Frank, Angel Hsing-Chi Hwang, Ruishan Liu
arXiv AI
Sep 15

ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information

ClinAgent is a conversational system that uses a ReAct-based LLM agent to retrieve and synthesize clinical trial information from multiple sources such as ClinicalTrials.gov, PubMed, and a local dataset. The agent iteratively reasons over user queries, selects appropriate tools, and refines its actions to provide grounded, up-to-date responses in natural language across multi-turn interactions. Evaluation across three phases shows that DeepSeek (thinking mode) excels in planning quality while Gemini 3.0 Flash delivers the highest overall performance and expert ratings, demonstrating the promise of agentic AI for improving clinical trial data access.

By Antonino Vaccarella, Riccardo Cantini, Domenico Talia, Paolo Trunfio, Marianna Talia, Rosamaria Lappano, Marcello Maggiolini