Hugging Face Trending Papers

RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems

arXiv AI
Sep 11

RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems

The paper introduces RESCUE-BENCH, a benchmark for relation-aware multi‑party emotional support conversation systems. It is built from real couple and family interview data, comprising 191 samples, 7,079 annotated turns, and 1,064.8 minutes of video, and defines six tasks that assess relational understanding and relation‑sensitive support. Experiments with ten large language models show that while they handle local emotional cues reasonably well, they struggle with tasks that require modeling interpersonal relations, such as predicting relation patterns, viewpoints, and support strategies.

By Haichuan Hu, Yang Xiao, Mingni Tang, Jiawen Duan, Quanjun Zhang, Congqing He, Hao Zhang, Jiashuo Wang, Johan F. Hoorn, Wenjie Li
arXiv AI
4d ago

From Momentary Emotion Inference to Sustained Emotion Support: Evaluating a Companion Agent in a Longitudinal Study

The study evaluates PAIR, a theory-based emotion‑regulation companion, over 14 days with 19 participants, analyzing 1,093 sessions. Emotion estimates from the agent aligned better with participants’ self‑reported valence and dominance than arousal, and guided conversations led to higher valence and state‑dependent arousal changes. Participants reported feeling understood, and the perceived helpfulness of guided conversations increased over time, highlighting the role of memory updates and cross‑session personalization in sustained emotional support.

By Kexin Quan, Zijian Ding, Jiaye Yong, Qinshi Zhang, Dong Wang, Jessie Chin
arXiv Computation and Language
Sep 3

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models

PIVOTSBench is a benchmark designed to assess multimodal large language models’ ability to reason about fine‑grained interpersonal relationships. It is constructed from Social‑IQ 2.0 and YouTube data and evaluates models on predicting bidirectional relationship dimensions grounded in psychology research. The benchmark also includes auxiliary tasks that test models’ capacity to identify and use critical visual cues, and it examines the impact of visual modalities, social role information, and different prediction settings on model performance.

By Shuxiang Zhang, Yiting Yin, Wenxuan Song, Yuhang Wu, Miao Liu
arXiv AI
Jul 17

From Stateless to Situated: Building a Psychological World for LLM-Based Agents

arXiv:2603. 25031v2 Announce Type: replace Abstract: In psychological support and emotional companionship scenarios, the core limitation of large language models (LLMs) lies not merely in response quality, but in their reliance on local next-token prediction, which prevents them from maintaining the temporal continuity, stage awareness, and user consent boundaries required for multi-turn intervention.

By Boning Zhao, Yutong Hu, Xinnuo Li
arXiv AI
Jul 10

From Triggers to Emotions: A CPM-Grounded Appraisal Multi-Agent for Dynamic Emotional Evolution in Persona-Based Dialogue

arXiv:2607. 07824v1 Announce Type: cross Abstract: Large Language Models (LLMs) have substantially advanced persona-based dialogue agents for emotion-sensitive role simulation in healthcare, education, counseling, customer service, and interactive storytelling.

By Jingyao Cai, Shuaijun Liu, Abdul Rehman, Yutong Guo, Qin Tian, Thomas Dolby, Sue Green, Chantel Cox, Xiaosong Yang
arXiv Computation and Language
5d ago

Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice

The paper introduces the Romantic Relationship Advice-Seeking Prompts (RRASP) dataset, comprising 2,400 prompts across five relationship themes, to study how query formulation affects sycophancy in large language models. Using the ELEPHANT framework, the authors evaluated GPT‑5 Mini and Gemini 3 Flash, finding that grammatical mood alone does not drive sycophantic behavior, whereas perspective‑driven framing does, with models increasingly accepting user premises over successive turns. Gemini 3 Flash showed smaller increases in moral sycophancy than GPT‑5 Mini, indicating greater resistance to reinforcing ethically problematic positions.

By Helena Choi, Edric Castel Hao, Karl Bautista, Francis Gabriel Magleo, Renzo Panti, Danielle Beatrice Olalia