Hugging Face Trending Papers

One Model, Multiple Goals: Adaptive Multi-Objective Learning for E-commerce Dialogue Systems

Dialogue systems in e-commerce scenarios often need to satisfy multiple objectives: accurately reasoning over user profiles (e. g.

arXiv AI
Sep 25

Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency

The paper introduces TRACER, a multi‑turn user simulator that models evolving user intent and aligns simulated behavior with real interaction trajectories. TRACER is trained first with supervised fine‑tuning on real dialogues and then with reinforcement learning that uses hierarchical outcome‑ and trajectory‑level rewards to address reward sparsity and credit assignment. In real customer‑service sessions, TRACER‑7B outperforms the best baseline by 11.4 conversion F1, achieves the lowest group‑level conversion‑rate error and semantic trajectory distance, and generalizes to out‑of‑distribution scenarios, while human Turing tests show its conversations appear natural. The authors also present the Dynamic Marketing Benchmark, which evaluates both persuasion effectiveness and response quality of large language models through simulated interactions, demonstrating that higher response quality does not always lead to higher conversion rates.

By Geng Chen, Ruotong Pan, Zhirui Yang, Qiqi He, Jiawei Chen, Zhang Yunfei, Chongyuan Chen, Minxuan Lv, Zheng Yang, Win-Bin Huang, Xiangyu Wu, Wenwu Ou
arXiv Computation and Language
Sep 4

RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents

RL-ADA introduces a co‑evolutionary training framework that replaces costly human annotations with world‑feedback rewards derived from interaction outcomes. In this system, a large Customer Support Agent and an Adversarial Customer Agent train together, guided by an automated judge that rewards successful resolution and realistic intent‑concealing utterances, respectively. Applied to a banking support proof of concept, the method eliminates routing errors and doubles the end‑to‑end PASS rate over five cycles, while also revealing a new adversarial strategy called Contextual Camouflage.

By Ram Narayanan, Harshit Rajgarhia, Abhishek Mukherji
arXiv Computation and Language
Sep 15

Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation

arXiv:2609.14648v1 Announce Type: new Abstract: Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often...

By Ziyi Zhu, Daniel R. Cahn, Thomas D. Hull, Caitlin A. Stamatis, Olivier Tieleman, Guilherme B. Freire, Jinghong Chen
arXiv Computation and Language
Sep 10

SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

SocialRL is a multi-turn reinforcement learning framework that refines large language models’ social intelligence. It uses PPO to propagate delayed outcome rewards across turns, enabling long‑horizon planning, and introduces six process reward dimensions—such as goal advancement and relational attunement—to capture the goal‑relationship trade‑off. A reward model provides fine‑grained scoring and a stage‑aware weight schedule prioritizes relationship building early, goal pursuit mid‑way, and balanced closure later, yielding an average 9.2 percentage‑point improvement in goal achievement across multiple social‑dialogue benchmarks.

By Jianing Wang, Xintao Wang, Aili Chen, Jie Shi, Hongcheng Guo, Jun Gao, Wenxuan Zhao, Chengkun Lang, Yuanli Guo, Yanghua Xiao
arXiv Computation and Language
Aug 25

ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents

arXiv:2608.21969v1 Announce Type: new Abstract: Humans have multiple levels of temporal abstractions on daily interaction and thinking, such as concept perception and strategic planning. Inspired by...

By Xiaoyu Wang, Qingqing Gu, Yue Zhao, Teng Chen, Yuqi Cao, Xiaokai Chen, Hongyan Li, Luo Ji
arXiv Machine Learning
Aug 27

Learning to summarize user information for personalized reinforcement learning from human feedback

The paper introduces PLUS, a framework that uses reinforcement learning to generate text-based summaries of individual users’ preferences, characteristics, and past conversations. These summaries condition a reward model, allowing it to predict personalized response preferences and improving reward accuracy by 11–77 % over the standard Bradley‑Terry model. PLUS demonstrates robust performance with new users and topics, achieves a 25 % improvement over existing personalized RLHF techniques, and enables zero‑shot personalization for state‑of‑the‑art models like GPT‑4.

By Hyunji Nam, Yanming Wan, Mickel Liu, Peter Ahnn, Jianxun Lian, Natasha Jaques
arXiv AI
Jun 15

UP-NRPA: User Portrait based Nested Rollout Policy Adaptation for Planning with Large Language Models in Goal-oriented Dialogue Systems

arXiv:2606. 13683v1 Announce Type: new Abstract: To address the challenge that current dialogue policy planning methods struggle to dynamically adapt to diverse user characteristics, this paper proposes a User Portrait based Nested Rollout Policy Adaptation (UP-NRPA) online framework with Large Language Models.

By Hui Wang, Fafa Zhang, Meng Liu, Xiangyu Chen, Chaoxu Mu
arXiv AI
Jul 7

Interactive Learning for LLM Reasoning

arXiv:2509. 26306v5 Announce Type: replace Abstract: Existing multi-agent learning approaches have developed interactive training environments to explicitly promote collaboration among multiple Large Language Models (LLMs), thereby constructing stronger multi-agent systems (MAS).

By Hehai Lin, Shilei Cao, Sudong Wang, Haotian Wu, Minzhi Li, Linyi Yang, Juepeng Zheng, Chengwei Qin