arXiv AI

Evaluation of Motivational Interviewing Counsellors with Task-Aware Multi-Stage LLM-Based Simulated Clients

arXiv:2608. 07499v1 Announce Type: cross Abstract: The development and benchmarking of Large Language Model (LLM)-based Motivational Interviewing (MI) counsellors now often rely on LLM-based simulated clients.

arXiv Computation and Language
Sep 18

Steering the Compass: Aligning Dynamic Psychological Counseling Conversations with Cognitive Behavioral Therapy Strategies

The paper introduces StratCBT, a new dataset of 9,688 psychological counseling sessions with 256K utterances, each counselor response aligned to one of eight Cognitive Behavioral Therapy (CBT) strategies. It was created by modeling clients’ negative thoughts and generating high‑quality conversations through self‑chat, using realistic sessions for guidance. Experiments show that strategy‑aligned generation improves professional and effective counseling when evaluated with large language model‑simulated clients.

By Zimu Wang, Yiwen Jiang, Xiangyu Zhao, Yaling Shen, Jiahe Liu, Stephanie Fong, Maxmartwell H Cheng, Guilherme C Oliveira, Anh Nguyen, Robert Desimone, Barnaby Nelson, Dominic Dwyer, Zongyuan Ge
arXiv Computation and Language
Sep 1

Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs

Graph2Counsel is a framework that generates synthetic counseling dialogues by leveraging Client Psychological Graphs (CPGs) to encode the relationships among a client’s thoughts, emotions, and behaviors. The system uses a structured prompting pipeline guided by counselor strategies and explores techniques such as Chain‑of‑Thought and Multi‑Agent Feedback to produce 760 realistic sessions from 76 CPGs. Expert evaluation shows the dataset surpasses previous ones in specificity, counselor competence, authenticity, conversational flow, and safety, and fine‑tuning an open‑source model on it improves performance on several counseling benchmarks.

By Aishik Mandal, Hiba Arnaout, Clarissa W. Ong, Juliet Bockhorst, Kate Sheehan, Rachael Moldow, Tanmoy Chakraborty, Iryna Gurevych
arXiv AI
Sep 10

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

The paper introduces a three-tier persona vector for user simulation in evaluating LLM agents, comprising 23 dimensions across demographics, behavioral traits, and emotional states, plus a query-complexity overlay. It demonstrates that these nuanced personas generate diverse, scenario-reactive conversations, leading to significant variations in agent goal achievement and compliance across different contexts. The model’s design allows for reproducible, auditable user behavior patterns without relying on learned covariance matrices.

By Rahul Khedar, Eshita, Sneha Teja Sree Reddy Thondapu, Mayank Malhotra, Arup Kumar Das, Jitesh Chandra Mishra, Arun Menon, Avinash Karn, Mouli V
Hugging Face Trending Papers
Sep 8

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

The paper introduces a three-tier persona vector to generate diverse, realistic user inputs for evaluating tool-augmented LLM agents. The vector includes 23 dimensions: categorical demographics, continuous behavioral traits, and continuous emotional states, plus a query-complexity overlay. Experiments on 64,698 conversations show that these persona dimensions produce measurable differences in agent performance and realistic scenario-reactive behavior.

arXiv Computation and Language
Aug 31

Roleplaying with Structure: Synthetic Therapist-Client Conversation Generation from Questionnaires

arXiv:2510.25384v2 Announce Type: replace Abstract: Large Language Models (LLMs) are promising tools for synthetic data generation in mental health. However, privacy policies and restrictions forced...

By Doan Nam Long Vu, Rui Tan, Lena Moench, Svenja Jule Francke, Daniel Woiwod, Florian Thomas-Odenthal, Sanna Stroth, Tilo Kircher, Christiane Hermann, Udo Dannlowski, Hamidreza Jamalabadi, Simone Balloccu, Shaoxiong Ji
arXiv AI
Jul 10

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

arXiv:2607. 08257v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters.

By Yuming Yang, Xiao Sun, Yuanwei Zou, Zhengxiao Wu, Yun Chen, Jiang Zhong, Haoyang Zeng, Jingwang Huang, Kaiwen Wei
arXiv Computation and Language
Sep 1

MedConceal: A Benchmark for Clinical Hidden-Concern Reasoning Under Partial Observability

MedConceal is a new benchmark for evaluating medical dialogue systems on hidden‑concern reasoning under partial observability. It features 300 curated cases and 600 clinician‑LLM interactions, using an interactive patient simulator that hides latent concerns and tracks their revelation and resolution through theory‑grounded communication signals. The benchmark assesses both confirmation (surfacing hidden concerns) and intervention (addressing the primary concern), revealing that current models excel on different metrics while human clinicians still outperform them on intervention success.

By Yikun Han, Joey Chan, Jingyuan Chen, Mengting Ai, Simo Du, Yue Guo