arXiv AI

Structured Prompting and Automated Evaluation in Fixed Synthetic Japanese-Language Counseling Dialogues

arXiv:2507. 02950v3 Announce Type: replace-cross Abstract: Large language models (LLMs) may support counseling training, yet evidence from Japanese-language interactions and automated quality ratings remains limited.

arXiv AI
Aug 28

Distinct Profiles of Run-to-Run Score Reliability and Expert-Panel Alignment Across Four LLM Evaluators of Simulated Japanese-Language AI-to-AI Counseling

The study examined how four large language models (GPT‑5.5, Gemini 3.5 Flash, Claude Opus 4.8, and Fable 5) scored 18 simulated Japanese‑language AI‑to‑AI counseling sessions compared to ratings from 15 human counseling experts. Each model evaluated every transcript three times on four motivational‑interviewing‑informed dimensions and overall quality, consistently giving higher scores for softening sustain talk and overall quality than the expert panel, though the magnitude varied by model. Run‑to‑run reliability (intraclass correlation coefficients ranging from .33 to .96) did not predict closer alignment with expert judgments, and the models’ ability to discriminate counselor conditions was distinct from both reliability and alignment.

By Keita Kiuchi, Yoshikazu Fujimoto, Hideyuki Got\=o, Tomonori Hosokawa, Makoto Nishimura, Y\=osuke Sat\=o, Izumi Sezai, Tomohiro Inoue
arXiv Computation and Language
Aug 28

Beyond Reflection: Affirmation as a Promising Behavioral Marker Associated with Quality in Text-Based Counseling

The paper investigates which counselor behaviors correlate with higher dialogue quality in AI-assisted text-based counseling. Using the KokoroChat dataset, the authors find that the strategy of affirmation consistently associates with better session quality, more so than reflection. Cross-dataset experiments suggest this signal also appears, to some extent, in an English dataset of non-expert supporters.

By Michimasa Inaba
arXiv Computation and Language
Sep 21

Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes

This scoping review examined 48 studies on generative AI chatbots designed to deliver motivational interviewing (MI). It found that most systems were text‑based and disembodied, with about half incorporating dynamic adaptation, and that safety reporting was inconsistent. While user perceptions were generally positive and many studies reported MI‑consistent interactions, evidence for sustained behavioral or functional change remains limited.

By Runze Hu, Jingqi Kong, Yang Yang, Yihang Yang, Jingyao Liu, Haizhou Tang, Shanghang Zhang, Zheng Liu
arXiv Computation and Language
Sep 1

Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs

Graph2Counsel is a framework that generates synthetic counseling dialogues by leveraging Client Psychological Graphs (CPGs) to encode the relationships among a client’s thoughts, emotions, and behaviors. The system uses a structured prompting pipeline guided by counselor strategies and explores techniques such as Chain‑of‑Thought and Multi‑Agent Feedback to produce 760 realistic sessions from 76 CPGs. Expert evaluation shows the dataset surpasses previous ones in specificity, counselor competence, authenticity, conversational flow, and safety, and fine‑tuning an open‑source model on it improves performance on several counseling benchmarks.

By Aishik Mandal, Hiba Arnaout, Clarissa W. Ong, Juliet Bockhorst, Kate Sheehan, Rachael Moldow, Tanmoy Chakraborty, Iryna Gurevych
arXiv Computation and Language
Sep 18

Steering the Compass: Aligning Dynamic Psychological Counseling Conversations with Cognitive Behavioral Therapy Strategies

The paper introduces StratCBT, a new dataset of 9,688 psychological counseling sessions with 256K utterances, each counselor response aligned to one of eight Cognitive Behavioral Therapy (CBT) strategies. It was created by modeling clients’ negative thoughts and generating high‑quality conversations through self‑chat, using realistic sessions for guidance. Experiments show that strategy‑aligned generation improves professional and effective counseling when evaluated with large language model‑simulated clients.

By Zimu Wang, Yiwen Jiang, Xiangyu Zhao, Yaling Shen, Jiahe Liu, Stephanie Fong, Maxmartwell H Cheng, Guilherme C Oliveira, Anh Nguyen, Robert Desimone, Barnaby Nelson, Dominic Dwyer, Zongyuan Ge
arXiv AI
Aug 24

When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

The paper evaluates the safety of conversational AI therapy bots for Generation Alpha, revealing that while these models understand 76‑82% of youth‑specific vocabulary, they correctly assess clinical risk only 64‑72% of the time, creating a significant vocabulary‑comprehension gap. Six failure patterns—such as sarcasm masking, minimization acceptance, and semantic drift—were identified, with compounded errors leading to a 94% miss rate when three or more patterns co‑occur. The authors estimate 146,880 missed crises annually and recommend mandatory human‑in‑the‑loop systems, quarterly youth‑specific validation, transparent performance disclosure, and regulatory oversight for youth‑facing mental health AI.

By Manisha Mehta, Virendra Mehta