arXiv:2606. 15532v1 Announce Type: cross Abstract: Emotional intelligence (EI) in Large Language Models (LLMs) is often evaluated through static understanding tasks or single-response dialogue generation.
By Rongzhi Zhu, Xiang Huang, Yuchuan Wu, Rui Wang, Zequn Sun, Tao Ren, Weiyao Luo, Bingxue Qiu, Jieping Ye, Yongbin Li, Wei Hu
The paper explores how dense, turn-level reward structures can improve reinforcement learning for large language model agents in multi-turn tasks. It introduces three reward granularity types—terminal, delayed, and per-turn—and adapts Group Relative Policy Optimization and Proximal Policy Optimization to each. Experiments on search and game agents show that per-turn rewards consistently yield better training dynamics, faster convergence, and higher answer correctness compared to sparse terminal or delayed rewards.
By Quan Wei, Siliang Zeng, Chenliang Li, Zhongruo Wang, William Brown, Oana Frunza, Wei Deng, Anderson Schneider, Yuriy Nevmyvaka, Yang Katie Zhao, Alfredo Garcia, Mingyi Hong
arXiv:2606. 11119v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models.
By Heming Zou, Qi Wang, Yun Qu, Yuhang Jiang, Lizhou Cai, Yixiu Mao, Ru Peng, Xin Xu, Weijie Liu, Kai Yang, Saiyong Yang, Xiangyang Ji
Large Language Model (LLM) agents increasingly solve long-horizon tasks through multi-turn interactions with users and external tools. In these settings, relevant task information often unfolds over t...
The paper introduces FACA, a Feedback‑Aware Credit Assignment method that aligns each agent reaction with the preceding user‑to‑user segment, derives a locally normalized reaction advantage, and adds it to the terminal outcome advantage without requiring an extra critic or rollout. Compared to an outcome‑only Interactive GRPO baseline, FACA improves performance across nine domains by 5.91–10.22 percentage points on 8B and 14B models, with notable gains in Telecom. The approach demonstrates that next‑turn user reactions provide actionable local credit for enhancing multi‑turn user‑interacting agents.
The paper introduces FACA, a Feedback‑Aware Credit Assignment method that aligns each agent reaction with the preceding user‑to‑user segment, computes a locally normalized reaction advantage, and adds it to the terminal outcome advantage without requiring an extra critic or rollout. Experiments show that FACA improves performance across nine domains, especially in Telecom, and maintains the same ordering of gains in zero‑shot benchmarks such as Pare‑Bench and Co‑Gym.
By Yiwen Zhao, Zhihao Wen, Yuchen Mao, Mingxuan Jiang, Yihao Hu, Pan Wang, Xin Zhang, Wei Wu
The paper introduces GRAFT, a Graph-based Faithful sTep-level credit-assignment framework that constructs a trajectory graph from rollout trajectories, recovers node state-values via Bellman iteration, and assigns step-level advantages based on node value differences. It also proposes Graph GAE to further reduce state-value estimation bias. Experiments on multi-turn agentic benchmarks demonstrate consistent improvements over GRPO and other recent agentic RL algorithms.
By Xincheng Yao, Haobo Fu, Weiming Liu, Chongyang Zhang
Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normal...
The paper introduces Potential-Guided Policy Optimization (PGPO), a method for multi-turn agentic tasks that improves credit assignment by estimating empirical state potentials from anchor-state-group return statistics. PGPO derives action advantages from potential differences between adjacent states, enabling cross-trajectory credit propagation and finer-grained step-level credit assignment, especially within failed trajectories. Experiments on ALFWorld and WebShop demonstrate strong performance compared to recent group-based reinforcement learning methods, with negligible training overhead.
By Yuyao Zheng, Haipeng Sun, Junwei Bao, Lemao Liu, Hongfei Jiang, Yang Song, Dejing Dou
arXiv:2608. 16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult.
By Huan Zhang, Mingju Chen, Dongxu Zhou, Can Lv, Heng Chang, Sen Cui, Faguo Wu, Shiji Zhou
arXiv:2607. 13988v1 Announce Type: new Abstract: Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundamental challenge during post-training.
By Leitian Tao, Baolin Peng, Wenlin Yao, Tao Ge, Hao Cheng, Mike Hang Wang, Jianfeng Gao, Sharon Li
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult. Existing approache...