Roleplaying with Structure: Synthetic Therapist-Client Conversation Generation from Questionnaires
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
Graph2Counsel is a framework that generates synthetic counseling dialogues by leveraging Client Psychological Graphs (CPGs) to encode the relationships among a client’s thoughts, emotions, and behaviors. The system uses a structured prompting pipeline guided by counselor strategies and explores techniques such as Chain‑of‑Thought and Multi‑Agent Feedback to produce 760 realistic sessions from 76 CPGs. Expert evaluation shows the dataset surpasses previous ones in specificity, counselor competence, authenticity, conversational flow, and safety, and fine‑tuning an open‑source model on it improves performance on several counseling benchmarks.
arXiv:2607. 24754v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and therapeutic appropriateness.
arXiv:2608.23248v1 Announce Type: cross Abstract: Traditional clinical prediction models rely on task-specific pipelines and curated, structured data, which scale poorly and underutilize unstructured...
arXiv:2509.04183v3 Announce Type: replace-cross Abstract: The growing demand for scalable psychological counseling highlights the need for high-quality, privacy-compliant data, yet such data remains...
The paper investigates the role of minimal responses—short, empathic utterances—in psychological counseling, noting that such brief replies are common in human dialogues but underrepresented in large language model (LLM) outputs. Using a two‑stage filtering approach and contextual verification with an LLM, the authors systematically analyze minimal responses across multiple counseling datasets. They find that while strong commercial LLMs can produce minimal replies when prompted, they often fail to judge when these replies are appropriate, and counseling‑specific models trained on synthetic data tend to generate longer, content‑rich responses instead.
HealthBench-Psych is a mental‑health subset of the OpenAI HealthBench benchmark, created by filtering 5,000 physician‑rubric conversations for mental‑health relevance using an LLM‑applied rubric and validating the selection through two rounds of blinded clinician review. The resulting 610 conversations (12.2 % of the corpus) are released along with a pipeline, model responses, grades, and analysis code. Evaluation of 20 frontier and open models by a cross‑vendor panel of three LLM judges shows a statistically tied frontier cluster, measurable refusal behavior in two models, and near‑identical rankings across judges (τ ≥ 0.92).