arXiv AI

PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning

arXiv:2508. 09521v3 Announce Type: replace-cross Abstract: Emotional support conversations require more than fluent responses.

arXiv AI
Jul 24

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

arXiv:2607. 21013v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion understanding with advanced video understanding abilities and natural language description.

By Lihuang Fang, Yuchen Zou, kebin Jin, Jinghui Qin
arXiv AI
6d ago

Affective Flow Language Model for Emotional Support Conversation

The paper introduces the Affective Flow Language Model (AFlow), which treats multi‑turn emotional support conversations as an evolving affective utility flow along dialogue trajectories. AFlow searches diverse support paths, estimates utilities of intermediate states, and employs Affective Flow Preference Optimization (AFPO) to propagate downstream preference signals to earlier states, enabling consistent strategy transitions. Experiments on ExTES and ESConv demonstrate improved strategy alignment, response diversity, and generation quality across various model settings.

By Chenghui Zou, Ning Wang, Tiesunlong Shen, Luwei Xiao, Chuan Ma, Xiangpeng Li, Rui Mao, Erik Cambria
arXiv AI
Jul 9

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

arXiv:2602. 23802v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture the complexity and subjectivity of human emotions.

By Yiyang Fang, Wenke Huang, Pei Fu, Yihao Yang, Kehua Su, Zhenbo Luo, Jian Luan, Mang Ye
arXiv Computation and Language
Sep 10

SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

SocialRL is a multi-turn reinforcement learning framework that refines large language models’ social intelligence. It uses PPO to propagate delayed outcome rewards across turns, enabling long‑horizon planning, and introduces six process reward dimensions—such as goal advancement and relational attunement—to capture the goal‑relationship trade‑off. A reward model provides fine‑grained scoring and a stage‑aware weight schedule prioritizes relationship building early, goal pursuit mid‑way, and balanced closure later, yielding an average 9.2 percentage‑point improvement in goal achievement across multiple social‑dialogue benchmarks.

By Jianing Wang, Xintao Wang, Aili Chen, Jie Shi, Hongcheng Guo, Jun Gao, Wenxuan Zhao, Chengkun Lang, Yuanli Guo, Yanghua Xiao
arXiv Computation and Language
Aug 25

ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents

arXiv:2608.21969v1 Announce Type: new Abstract: Humans have multiple levels of temporal abstractions on daily interaction and thinking, such as concept perception and strategic planning. Inspired by...

By Xiaoyu Wang, Qingqing Gu, Yue Zhao, Teng Chen, Yuqi Cao, Xiaokai Chen, Hongyan Li, Luo Ji