Graph2Counsel is a framework that generates synthetic counseling dialogues by leveraging Client Psychological Graphs (CPGs) to encode the relationships among a client’s thoughts, emotions, and behaviors. The system uses a structured prompting pipeline guided by counselor strategies and explores techniques such as Chain‑of‑Thought and Multi‑Agent Feedback to produce 760 realistic sessions from 76 CPGs. Expert evaluation shows the dataset surpasses previous ones in specificity, counselor competence, authenticity, conversational flow, and safety, and fine‑tuning an open‑source model on it improves performance on several counseling benchmarks.
By Aishik Mandal, Hiba Arnaout, Clarissa W. Ong, Juliet Bockhorst, Kate Sheehan, Rachael Moldow, Tanmoy Chakraborty, Iryna Gurevych
arXiv:2607. 24754v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and therapeutic appropriateness.
By Asher Sprigler, Yixue Zhao, Yi Ding
arXiv:2608.23248v1 Announce Type: cross
Abstract: Traditional clinical prediction models rely on task-specific pipelines and curated, structured data, which scale poorly and underutilize unstructured...
By Siri Willems, James Butterworth, Lore Goetschalckx, Peter Vrancx, Philippe Modard, Elke Giets, Ludovic Denoyer
arXiv:2509.04183v3 Announce Type: replace-cross
Abstract: The growing demand for scalable psychological counseling highlights the need for high-quality, privacy-compliant data, yet such data remains...
By Aishik Mandal, Tanmoy Chakraborty, Iryna Gurevych
The paper investigates the role of minimal responses—short, empathic utterances—in psychological counseling, noting that such brief replies are common in human dialogues but underrepresented in large language model (LLM) outputs. Using a two‑stage filtering approach and contextual verification with an LLM, the authors systematically analyze minimal responses across multiple counseling datasets. They find that while strong commercial LLMs can produce minimal replies when prompted, they often fail to judge when these replies are appropriate, and counseling‑specific models trained on synthetic data tend to generate longer, content‑rich responses instead.
By Zhiyang Qi
HealthBench-Psych is a mental‑health subset of the OpenAI HealthBench benchmark, created by filtering 5,000 physician‑rubric conversations for mental‑health relevance using an LLM‑applied rubric and validating the selection through two rounds of blinded clinician review. The resulting 610 conversations (12.2 % of the corpus) are released along with a pipeline, model responses, grades, and analysis code. Evaluation of 20 frontier and open models by a cross‑vendor panel of three LLM judges shows a statistically tied frontier cluster, measurable refusal behavior in two models, and near‑identical rankings across judges (τ ≥ 0.92).
By Matthew Flathers, Phuong Anh Nguyen, Jill Noorily, Julian Herpertz, Meiting Chen, Jasreen Multani, Samuel Powell, Mason Granof, Mark Kalinch, John Torous
arXiv:2607. 08257v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters.
By Yuming Yang, Xiao Sun, Yuanwei Zou, Zhengxiao Wu, Yun Chen, Jiang Zhong, Haoyang Zeng, Jingwang Huang, Kaiwen Wei
arXiv:2606. 19640v1 Announce Type: cross Abstract: AI and large language models (LLMs) have emerged as promising tools to address global mental health challenges.
By Yunkai Xu, Saeed Abdullah
The paper reviews how large language models are applied in mental health, covering areas such as social media analysis, clinical conversational agents, therapy support tools, prompt engineering, and multimodal learning. It synthesizes interdisciplinary studies that use social media posts, electronic medical records, and multimodal inputs to detect depression, assess suicide risk, provide personalized therapy, and generate psychoeducational content. The review also discusses advances in model interpretability, annotation strategies, multimodal fusion techniques, and highlights ethical, sociotechnical, and regulatory challenges while proposing frameworks for safe, equitable, and accountable deployment.
By Yisong Chen, Yifan Gao, Sijing Yu, Chuqing Zhao, Yang Lu
The paper evaluates the safety of conversational AI therapy bots for Generation Alpha, revealing that while these models understand 76‑82% of youth‑specific vocabulary, they correctly assess clinical risk only 64‑72% of the time, creating a significant vocabulary‑comprehension gap. Six failure patterns—such as sarcasm masking, minimization acceptance, and semantic drift—were identified, with compounded errors leading to a 94% miss rate when three or more patterns co‑occur. The authors estimate 146,880 missed crises annually and recommend mandatory human‑in‑the‑loop systems, quarterly youth‑specific validation, transparent performance disclosure, and regulatory oversight for youth‑facing mental health AI.
By Manisha Mehta, Virendra Mehta
DocTalkBN is a large-scale multimodal dataset of authentic expert telemedicine conversations in Bengali, comprising 557.63 hours of paired audio and text, 1,515 multi-turn patient calls, and 10,274 host–doctor question–answer exchanges across 26 medical specialties. The dataset contains 1.7 million tokens and preserves the spontaneity and contextual richness of real medical interactions in a low-resource language. Three downstream tasks—medical triage classification, advice safety evaluation, and medical named entity recognition—are constructed to benchmark large language models and encoder-based baselines, demonstrating DocTalkBN’s practical usefulness for clinically grounded reasoning.
By Anik Saha, Fahmida Sultana Naznin, Sadatul Islam Sadi, Ananya Shahrin Promi, Wahid Al Azad Navid, Rifat Shahriyar
arXiv:2606. 26879v1 Announce Type: new Abstract: Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is restricted.
By William Poulett