arXiv Machine Learning

StagePilot: Stage-Level Planning for Long-Horizon Dialogue Simulation in Cybergrooming

arXiv:2602. 05060v2 Announce Type: replace Abstract: Cybergrooming is an evolving threat to youth, requiring proactive educational interventions.

arXiv Computation and Language
Sep 10

SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

SocialRL is a multi-turn reinforcement learning framework that refines large language models’ social intelligence. It uses PPO to propagate delayed outcome rewards across turns, enabling long‑horizon planning, and introduces six process reward dimensions—such as goal advancement and relational attunement—to capture the goal‑relationship trade‑off. A reward model provides fine‑grained scoring and a stage‑aware weight schedule prioritizes relationship building early, goal pursuit mid‑way, and balanced closure later, yielding an average 9.2 percentage‑point improvement in goal achievement across multiple social‑dialogue benchmarks.

By Jianing Wang, Xintao Wang, Aili Chen, Jie Shi, Hongcheng Guo, Jun Gao, Wenxuan Zhao, Chengkun Lang, Yuanli Guo, Yanghua Xiao
arXiv Computation and Language
Sep 1

SIC-Agents: Benchmarking and Building an Adaptive Simulator for Pediatric Serious Illness Communication Training

The paper introduces SIC-Agents, a self‑improving framework designed to enhance simulation for pediatric serious illness communication (SIC) training. It presents two new benchmark suites—PitfallBench and DialogueBench—that assess simulators at both turn‑level and full‑dialogue levels, specifically addressing the unique challenges of multi‑party interactions and parental distress. Experiments demonstrate that SIC‑Agents surpasses static expert prompting, and the authors release the benchmarks for broader research use.

By Zihan Wang, Anita Marie Slominska, Rennie Bimman, Elizabeth Di Flumeri, Amanda Mayappo-Neeposh, Conall Francoeur, Tamara Ellen Carver, Xiao-Wen Chang, Doina Precup, Esin Darici Haritaoglu, Ismail Haritaoglu, Akshatha Arodi, Naomi Goloff
arXiv AI
Jul 10

SimRPD: Optimizing Recruitment Proactive Dialogue Agents through Simulator-Based Data Evaluation and Selection

arXiv:2601. 02871v3 Announce Type: replace Abstract: Task-oriented proactive dialogue agents play a pivotal role in recruitment, particularly for steering conversations towards specific business outcomes, such as acquiring social-media contacts for private-channel conversion.

By Zhiyong Cao, Dunqiang Liu, Qi Dai, Haojun Xu, Huai Yuen Khor, Hao Wang, Huan He, Yafei Liu, Ke Ma, Ruqian Shi, Sicheng Zhou, Sijia Yao
arXiv AI
Jun 15

UP-NRPA: User Portrait based Nested Rollout Policy Adaptation for Planning with Large Language Models in Goal-oriented Dialogue Systems

arXiv:2606. 13683v1 Announce Type: new Abstract: To address the challenge that current dialogue policy planning methods struggle to dynamically adapt to diverse user characteristics, this paper proposes a User Portrait based Nested Rollout Policy Adaptation (UP-NRPA) online framework with Large Language Models.

By Hui Wang, Fafa Zhang, Meng Liu, Xiangyu Chen, Chaoxu Mu
arXiv AI
Jul 21

Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment

arXiv:2607. 17191v1 Announce Type: new Abstract: Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc.

By Wentao Liu, Siyu Song, Xi Chen, Youjia Li, Xiaokun Wang, Min Ji, Ji Wang
arXiv Computation and Language
Aug 25

ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents

arXiv:2608.21969v1 Announce Type: new Abstract: Humans have multiple levels of temporal abstractions on daily interaction and thinking, such as concept perception and strategic planning. Inspired by...

By Xiaoyu Wang, Qingqing Gu, Yue Zhao, Teng Chen, Yuqi Cao, Xiaokai Chen, Hongyan Li, Luo Ji