PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning
arXiv:2606. 00395v1 Announce Type: cross Abstract: Mixture of Experts (MoE) Large Language Models (LLMs) achieve strong performance at scale.
arXiv:2606. 12479v1 Announce Type: cross Abstract: Large language model (LLM) routing has emerged as an effective paradigm for leveraging the complementary strengths of multiple LLMs through dynamic model and reasoning-strategy selection.
arXiv:2606. 00395v1 Announce Type: cross Abstract: Mixture of Experts (MoE) Large Language Models (LLMs) achieve strong performance at scale.
arXiv:2606. 15866v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective post-training paradigm for improving the reasoning abilities of large language models.
arXiv:2601. 06487v3 Announce Type: replace-cross Abstract: Reinforcement learning has substantially improved the performance of LLM agents on tasks with verifiable outcomes, but it still struggles on open-ended agent tasks with vast solution spaces (e.
arXiv:2507. 04136v2 Announce Type: replace Abstract: This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Proximal Policy Optimization (PPO), Q-Learning, and Actor-Critic methods.
ArenaFlow is a hierarchical credit propagation framework designed to improve reinforcement learning for open-ended agent tasks. It uses tournament-based relative ranking to generate trajectory-level rewards and structured reflective evaluation to identify pivotal success steps, reusable strategy skills, and skill usage attribution. The framework propagates advantages to high-confidence steps and maintains a global skill memory, enabling more targeted optimization and reusable skill priors for future exploration.
arXiv:2609.13058v1 Announce Type: new Abstract: Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models hav...
CounterRoute is an online reinforcement‑learning framework that jointly learns how to route a language model’s reasoning between a ‘think’ and a ‘direct answer’ mode, using a single shared policy derived from a dual‑mode checkpoint. It employs counterfactual rollouts to credit routing decisions and a curriculum that starts with forced dual‑mode rollouts before shifting to self‑routed updates, achieving better accuracy‑efficiency trade‑offs across nine benchmarks. The method reduces generated tokens by up to 51% on Qwen3‑8B while improving macro‑average accuracy, and its routing strategy generalizes to unseen coding, science, and commonsense tasks.
arXiv:2607. 04713v1 Announce Type: cross Abstract: Reinforcement learning holds significant potential for training large language models (LLMs) to handle multi-turn interactive tasks.
The paper explores how dense, turn-level reward structures can improve reinforcement learning for large language model agents in multi-turn tasks. It introduces three reward granularity types—terminal, delayed, and per-turn—and adapts Group Relative Policy Optimization and Proximal Policy Optimization to each. Experiments on search and game agents show that per-turn rewards consistently yield better training dynamics, faster convergence, and higher answer correctness compared to sparse terminal or delayed rewards.
arXiv:2604. 02621v2 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) substantially improves the reasoning capabilities of language models, but most existing RL fine-tuning approaches rely entirely on ground-truth verifiable rewards and thus labeled datasets with verifiable answers.
arXiv:2607. 06987v1 Announce Type: new Abstract: Reinforcement learning (RL) has become the standard paradigm for enhancing the complex reasoning capabilities of large language models (LLMs).
The paper introduces MEMENTO, a memory‑enhanced neural solver that improves routing problem solutions by using online data from repeated attempts to adjust action distributions during inference. It targets NP‑hard routing tasks such as the Traveling Salesman and Capacitated Vehicle Routing problems, outperforming existing tree‑search and policy‑gradient fine‑tuning methods. MEMENTO demonstrates strong scalability and data efficiency, achieving state‑of‑the‑art results on 11 of 12 evaluated tasks and enabling zero‑shot integration with diversity‑based solvers.