arXiv Machine Learning

EIBench: A Simulator-Based Benchmark and Turn-Credit RL for Emotion Management

arXiv:2606. 15532v1 Announce Type: cross Abstract: Emotional intelligence (EI) in Large Language Models (LLMs) is often evaluated through static understanding tasks or single-response dialogue generation.

arXiv AI
Jul 31

MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue

arXiv:2603. 06194v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for large language models (LLMs) has shown strong performance in single-turn tasks, but extending it to multi-turn interaction remains challenging due to sparse rewards and poor per-turn credit assignment.

By Naifan Zhang, Ruihan Sun, Jinwei Su, Hengjie Yang, Zhengyuan Pan, Zhaohan Chen, Xiaofan Zhang
arXiv AI
Jun 3

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments

arXiv:2606. 03892v1 Announce Type: cross Abstract: Training LLMs to orchestrate multi-step tool calls is held back by three coupled obstacles: realistic stateful execution environments are costly to build, synthetic training queries are often detached from the server's actual state (so the generated tool calls fail to execute), and recall-based RL rewards incentivize verbose tool-calling patterns.

By Ibrahim Abdelaziz, Asim Munawar, Kinjal Basu, Maxwell Crouse, Chulaka Gunasekara, Suneet Katrekar, Pavan Kapanipathi
arXiv AI
Jun 10

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

arXiv:2606. 11119v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models.

By Heming Zou, Qi Wang, Yun Qu, Yuhang Jiang, Lizhou Cai, Yixiu Mao, Ru Peng, Xin Xu, Weijie Liu, Kai Yang, Saiyong Yang, Xiangyang Ji
arXiv AI
3d ago

FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

The paper introduces FIGS, a dual‑axis evaluation framework for multi‑turn sycophancy that avoids penalizing empathy. It uses a 10‑turn conversational simulator with 500 diverse scenarios to test whether models stay truthful while keeping praise proportional, and whether they show calibrated validation of user feelings. The study finds that current models either drift toward sycophancy or become overly detached, highlighting an unresolved trade‑off in sustained dialogue.

By Sidharth Pulipaka, Ruta Binkyte, Ivaxi Sheth, Sahar Abdelnabi
arXiv AI
Jul 21

Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment

arXiv:2607. 17191v1 Announce Type: new Abstract: Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc.

By Wentao Liu, Siyu Song, Xi Chen, Youjia Li, Xiaokun Wang, Min Ji, Ji Wang
arXiv Machine Learning
Sep 2

Iterative GRPO: Batch-Online Policy Iteration for Multi-Turn RL via Single-Turn RLHF

Iterative GRPO is a batch‑online policy iteration framework that enables multi‑turn reinforcement learning for conversational agents without requiring an interactive user simulator. It alternates between learning a turn‑level Q‑function from logged returns (policy evaluation) and applying single‑turn GRPO against this Q‑function (policy improvement), thereby scoring candidate responses by their expected downstream return. The method is validated on six multi‑turn negotiation environments, demonstrating its practicality for real‑world deployment patterns.

By Daniel R. Jiang, Ankur Samanta, Yukai Yang, Jalaj Bhandari, R\'emi Munos, Tyler Lu
arXiv AI
Aug 19

Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context

The paper introduces FACA, a Feedback‑Aware Credit Assignment method that aligns each agent reaction with the preceding user‑to‑user segment, computes a locally normalized reaction advantage, and adds it to the terminal outcome advantage without requiring an extra critic or rollout. Experiments show that FACA improves performance across nine domains, especially in Telecom, and maintains the same ordering of gains in zero‑shot benchmarks such as Pare‑Bench and Co‑Gym.

By Yiwen Zhao, Zhihao Wen, Yuchen Mao, Mingxuan Jiang, Yihao Hu, Pan Wang, Xin Zhang, Wei Wu
Hugging Face Trending Papers
Aug 18

Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context

The paper introduces FACA, a Feedback‑Aware Credit Assignment method that aligns each agent reaction with the preceding user‑to‑user segment, derives a locally normalized reaction advantage, and adds it to the terminal outcome advantage without requiring an extra critic or rollout. Compared to an outcome‑only Interactive GRPO baseline, FACA improves performance across nine domains by 5.91–10.22 percentage points on 8B and 14B models, with notable gains in Telecom. The approach demonstrates that next‑turn user reactions provide actionable local credit for enhancing multi‑turn user‑interacting agents.