RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents
arXiv:2607. 04713v1 Announce Type: cross Abstract: Reinforcement learning holds significant potential for training large language models (LLMs) to handle multi-turn interactive tasks.
arXiv:2606. 24133v1 Announce Type: new Abstract: The composition of training data, governed by the diversity of sources and their mixing strategy, is a cornerstone of Large Language Model (LLM) pre-training.
arXiv:2607. 04713v1 Announce Type: cross Abstract: Reinforcement learning holds significant potential for training large language models (LLMs) to handle multi-turn interactive tasks.
arXiv:2606. 07705v1 Announce Type: cross Abstract: Although multi-objective reinforcement learning (MORL) is central to aligning large language models with complex human preferences, the prevailing practice of static weighted summation overlooks a more fundamental phenomenon: reward learning is markedly asynchronous across objectives.
EmbodiedMind introduces a three-stage training paradigm for embodied foundation models that tackles inefficient sample use, task imbalance, and credit assignment in long-horizon planning. The stages—Rejection Sampling-based Fine‑Tuning, Iterative Rejection GRPO, and Trie‑GRPO—filter low‑informative data, balance task difficulty, and use action prefix trees for step‑level advantage estimation. This approach yields a state‑of‑the‑art average performance of 70.02% across 18 benchmarks, notably improving long‑horizon task planning accuracy.
arXiv:2608.20909v1 Announce Type: new Abstract: Offline RL methods commonly jointly train the actor and critic, where the critic is used to guide the actor toward higher-value actions. This coupled l...
The paper introduces Bidirectional Behavior Prior Distillation (B2PD), a method that uses action‑value priors to train a conditional variational autoencoder for generating high‑value behavior support. These expert behavior priors are then distilled into the online reinforcement learning agent, reducing inefficient exploration and stabilizing policy updates. Experiments on state‑ and pixel‑based tasks show that B2PD improves sample efficiency while maintaining stable learning dynamics.
arXiv:2606. 06673v1 Announce Type: new Abstract: Sparse rewards and heterogeneous task sequences remain persistent challenges in Reinforcement Learning (RL), often resulting in slow convergence, weak generalization, and inefficient exploration.
arXiv:2608. 19684v1 Announce Type: new Abstract: Recent studies investigate how to leverage pre-collected datasets to improve the policy performance and sample efficiency of RL.
arXiv:2603.26164v2 Announce Type: replace-cross Abstract: Data-centric training has emerged as a promising direction for improving large language models (LLMs) by optimizing not only model parameters...
arXiv:2502. 19544v3 Announce Type: replace Abstract: Leveraging offline data is a promising way to improve the sample efficiency of online reinforcement learning (RL).
arXiv:2607. 27973v1 Announce Type: new Abstract: Recently, Reinforcement Learning (RL) has emerged as a crucial paradigm for the post-training of Large Language Model (LLM) agents.
arXiv:2609.00213v1 Announce Type: new Abstract: Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific e...
The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.