While AI-assisted text-based counseling is gaining attention, it remains empirically unclear which counselor behaviors are associated with higher dialogue quality. Existing research often focuses heav...
arXiv:2507. 02950v3 Announce Type: replace-cross Abstract: Large language models (LLMs) may support counseling training, yet evidence from Japanese-language interactions and automated quality ratings remains limited.
By Keita Kiuchi, Yoshikazu Fujimoto, Hideyuki Goto, Tomonori Hosokawa, Makoto Nishimura, Yosuke Sato, Izumi Sezai, Tomohiro Inoue
The paper investigates the role of minimal responses—short, empathic utterances—in psychological counseling, noting that such brief replies are common in human dialogues but underrepresented in large language model (LLM) outputs. Using a two‑stage filtering approach and contextual verification with an LLM, the authors systematically analyze minimal responses across multiple counseling datasets. They find that while strong commercial LLMs can produce minimal replies when prompted, they often fail to judge when these replies are appropriate, and counseling‑specific models trained on synthetic data tend to generate longer, content‑rich responses instead.
By Zhiyang Qi
The study examined how four large language models (GPT‑5.5, Gemini 3.5 Flash, Claude Opus 4.8, and Fable 5) scored 18 simulated Japanese‑language AI‑to‑AI counseling sessions compared to ratings from 15 human counseling experts. Each model evaluated every transcript three times on four motivational‑interviewing‑informed dimensions and overall quality, consistently giving higher scores for softening sustain talk and overall quality than the expert panel, though the magnitude varied by model. Run‑to‑run reliability (intraclass correlation coefficients ranging from .33 to .96) did not predict closer alignment with expert judgments, and the models’ ability to discriminate counselor conditions was distinct from both reliability and alignment.
By Keita Kiuchi, Yoshikazu Fujimoto, Hideyuki Got\=o, Tomonori Hosokawa, Makoto Nishimura, Y\=osuke Sat\=o, Izumi Sezai, Tomohiro Inoue
arXiv:2608.22251v1 Announce Type: cross
Abstract: Rising global demand for mental health support creates significant service delivery challenges, with asynchronous email counselling serving as a cruc...
By Philipp Steigerwald, Nico Bienlein, Jennifer Burghardt, Mara Stieler, Robert Lehmann, Jens Albrecht
Graph2Counsel is a framework that generates synthetic counseling dialogues by leveraging Client Psychological Graphs (CPGs) to encode the relationships among a client’s thoughts, emotions, and behaviors. The system uses a structured prompting pipeline guided by counselor strategies and explores techniques such as Chain‑of‑Thought and Multi‑Agent Feedback to produce 760 realistic sessions from 76 CPGs. Expert evaluation shows the dataset surpasses previous ones in specificity, counselor competence, authenticity, conversational flow, and safety, and fine‑tuning an open‑source model on it improves performance on several counseling benchmarks.
By Aishik Mandal, Hiba Arnaout, Clarissa W. Ong, Juliet Bockhorst, Kate Sheehan, Rachael Moldow, Tanmoy Chakraborty, Iryna Gurevych
arXiv:2607.23648v2 Announce Type: replace
Abstract: Using large language models (LLMs) to assist psychological counseling is an important task in the field of natural language processing. The constru...
By Kaitong Weng, Lixin Liu, Zihao Liu, Bo Wang, Shiguang Ni
arXiv:2608.31007v1 Announce Type: cross
Abstract: Understanding how psychiatric patients subjectively experienced a clinical conversation is important for feedback and alliance-related process monito...
By Aowen Shi, Michal Balazia, Danilo Postin, Ren\'e Hurlemann, Jan Alexandersson, Fran\c{c}ois Br\'emond, Philipp M\"uller
arXiv:2608. 02046v2 Announce Type: replace-cross Abstract: LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated.
By Yao Liu, Guangjia Chai, Yuming Huang, Jihao Huang, Lei Wang, Junchen Wan
LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same-family favoritism and scale drift.
arXiv:2608.29326v1 Announce Type: cross
Abstract: Positive psychology dialogue aims to support emotional distress and positive resource building, requiring models to produce not only empathetic repli...
By Yuxiong Wang, Ziwei Lin, Bo Wang, Yu Zhang, Shiguang Ni
arXiv:2608. 07499v1 Announce Type: cross Abstract: The development and benchmarking of Large Language Model (LLM)-based Motivational Interviewing (MI) counsellors now often rely on LLM-based simulated clients.
By Jiading Zhu, Xinyu Cindy Wang, Thomas Nguyen, Yan Qing Lee, Osnat C. Melamed, Peter Selby, Jonathan Rose