arXiv AI By Mohit Chandra, Nabin Kim, Eli Min, Aamogh Sawant, Tanmay Sutar, Munmun De Choudhury

Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation

Read the original on arXiv AI →

The paper introduces the COmmunity-centered Peer Engaged Support (COPES) dataset and a three‑axis evaluation framework to gauge how well Large Language Models (LLMs) align with community perspectives on mental‑health support queries. Experiments show that fine‑tuning LLMs on COPES improves strategy alignment and emotion‑tone alignment by over 50% for general‑purpose models, yet these gains are uneven across subreddits and coping strategies. The study also finds that post‑training shifts the model’s recommendations toward problem‑focused advice while reducing emotion‑focused responses, indicating persistent disparities in performance across different communities and needs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs

Graph2Counsel is a framework that generates synthetic counseling dialogues by leveraging Client Psychological Graphs (CPGs) to encode the relationships among a client’s thoughts, emotions, and behaviors. The system uses a structured prompting pipeline guided by counselor strategies and explores techniques such as Chain‑of‑Thought and Multi‑Agent Feedback to produce 760 realistic sessions from 76 CPGs. Expert evaluation shows the dataset surpasses previous ones in specificity, counselor competence, authenticity, conversational flow, and safety, and fine‑tuning an open‑source model on it improves performance on several counseling benchmarks.

By Aishik Mandal, Hiba Arnaout, Clarissa W. Ong, Juliet Bockhorst, Kate Sheehan, Rachael Moldow, Tanmoy Chakraborty, Iryna Gurevych
arXiv Computation and Language
Sep 1

Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts

The study investigates how large language models (LLMs) assess psychological distress in online posts from six identity‑based communities. Through a perspectivist annotation task, 321 participants provided 9,587 judgments on 1,198 Reddit posts, revealing modest in‑group agreement (OR = 1.18) that varies across communities. When evaluated against these community‑specific labels, open‑weight LLMs consistently over‑estimate distress—achieving only 31–44% accuracy on posts perceived as none‑to‑mild—while newer models like GPT‑5 and Gemini 2.5 Pro show similar inflation, whereas Claude Opus 4 is more conservative. "whyItMatters":"The findings highlight that miscalibrated distress detection by LLMs can disproportionately impact the very communities they aim to serve, underscoring the need for equitable AI deployment in mental‑health contexts."

By Andrew Aquilina, Xiang Lorraine Li, Yu-Ru Li
arXiv AI
Sep 11

RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems

The paper introduces RESCUE-BENCH, a benchmark for relation-aware multi‑party emotional support conversation systems. It is built from real couple and family interview data, comprising 191 samples, 7,079 annotated turns, and 1,064.8 minutes of video, and defines six tasks that assess relational understanding and relation‑sensitive support. Experiments with ten large language models show that while they handle local emotional cues reasonably well, they struggle with tasks that require modeling interpersonal relations, such as predicting relation patterns, viewpoints, and support strategies.

By Haichuan Hu, Yang Xiao, Mingni Tang, Jiawen Duan, Quanjun Zhang, Congqing He, Hao Zhang, Jiashuo Wang, Johan F. Hoorn, Wenjie Li