arXiv:2603. 06194v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for large language models (LLMs) has shown strong performance in single-turn tasks, but extending it to multi-turn interaction remains challenging due to sparse rewards and poor per-turn credit assignment.
By Naifan Zhang, Ruihan Sun, Jinwei Su, Hengjie Yang, Zhengyuan Pan, Zhaohan Chen, Xiaofan Zhang
arXiv:2606. 03892v1 Announce Type: cross Abstract: Training LLMs to orchestrate multi-step tool calls is held back by three coupled obstacles: realistic stateful execution environments are costly to build, synthetic training queries are often detached from the server's actual state (so the generated tool calls fail to execute), and recall-based RL rewards incentivize verbose tool-calling patterns.
By Ibrahim Abdelaziz, Asim Munawar, Kinjal Basu, Maxwell Crouse, Chulaka Gunasekara, Suneet Katrekar, Pavan Kapanipathi
arXiv:2606. 14199v1 Announce Type: cross Abstract: Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation.
By Xuhui Zhou, Weiwei Sun, Weihua Du, Jiarui Liu, Haojia Sun, Qianou Ma, Tongshuang Wu, Yiming Yang, Maarten Sap
arXiv:2607. 07508v1 Announce Type: cross Abstract: Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs).
By Zhenyu Hou, Yujiang Li, Jie Tang, Yuxiao Dong
arXiv:2606. 11119v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models.
By Heming Zou, Qi Wang, Yun Qu, Yuhang Jiang, Lizhou Cai, Yixiu Mao, Ru Peng, Xin Xu, Weijie Liu, Kai Yang, Saiyong Yang, Xiangyang Ji
arXiv:2608. 09168v1 Announce Type: new Abstract: Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge.
By Liang He, Jingbo Wen, Hongyu Gu, Hao Li, Haoyu Wang, Yixiong Chen, Kangning Cui, Xilu Wang
arXiv:2607. 27816v2 Announce Type: replace-cross Abstract: Role-playing agents (RPAs) have become one of the most important consumer applications of large language models.
By Yuhang Zhu, Mingxuan Du, Benfeng Xu, Jie Gao, Lingyun Yu, Hongtao Xie
The paper introduces FIGS, a dual‑axis evaluation framework for multi‑turn sycophancy that avoids penalizing empathy. It uses a 10‑turn conversational simulator with 500 diverse scenarios to test whether models stay truthful while keeping praise proportional, and whether they show calibrated validation of user feelings. The study finds that current models either drift toward sycophancy or become overly detached, highlighting an unresolved trade‑off in sustained dialogue.
By Sidharth Pulipaka, Ruta Binkyte, Ivaxi Sheth, Sahar Abdelnabi
arXiv:2607. 17191v1 Announce Type: new Abstract: Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc.
By Wentao Liu, Siyu Song, Xi Chen, Youjia Li, Xiaokun Wang, Min Ji, Ji Wang
Iterative GRPO is a batch‑online policy iteration framework that enables multi‑turn reinforcement learning for conversational agents without requiring an interactive user simulator. It alternates between learning a turn‑level Q‑function from logged returns (policy evaluation) and applying single‑turn GRPO against this Q‑function (policy improvement), thereby scoring candidate responses by their expected downstream return. The method is validated on six multi‑turn negotiation environments, demonstrating its practicality for real‑world deployment patterns.
By Daniel R. Jiang, Ankur Samanta, Yukai Yang, Jalaj Bhandari, R\'emi Munos, Tyler Lu
The paper introduces FACA, a Feedback‑Aware Credit Assignment method that aligns each agent reaction with the preceding user‑to‑user segment, computes a locally normalized reaction advantage, and adds it to the terminal outcome advantage without requiring an extra critic or rollout. Experiments show that FACA improves performance across nine domains, especially in Telecom, and maintains the same ordering of gains in zero‑shot benchmarks such as Pare‑Bench and Co‑Gym.
By Yiwen Zhao, Zhihao Wen, Yuchen Mao, Mingxuan Jiang, Yihao Hu, Pan Wang, Xin Zhang, Wei Wu
The paper introduces FACA, a Feedback‑Aware Credit Assignment method that aligns each agent reaction with the preceding user‑to‑user segment, derives a locally normalized reaction advantage, and adds it to the terminal outcome advantage without requiring an extra critic or rollout. Compared to an outcome‑only Interactive GRPO baseline, FACA improves performance across nine domains by 5.91–10.22 percentage points on 8B and 14B models, with notable gains in Telecom. The approach demonstrates that next‑turn user reactions provide actionable local credit for enhancing multi‑turn user‑interacting agents.