arXiv:2609.13058v1 Announce Type: new
Abstract: Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models hav...
By Hongyi He, Zhenghao Lin, Xiao Liu, Peng Cheng, Yan Lu, Yeyun Gong
arXiv:2606. 12479v1 Announce Type: cross Abstract: Large language model (LLM) routing has emerged as an effective paradigm for leveraging the complementary strengths of multiple LLMs through dynamic model and reasoning-strategy selection.
By Qihang Yu, Hanwen Tong, Zhengqi Zhang, Bo Zheng, Feng Wei, Shengyu Zhang, Zemin Liu, Fei Wu
The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.
By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng
arXiv:2609.08232v1 Announce Type: cross
Abstract: Detailed routing remains a dominant runtime bottleneck in physical design due to increasing complexity of design rules. Modern routers can struggle t...
By Afsara Khan, Austin Rovinski
arXiv:2608. 12146v1 Announce Type: cross Abstract: Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems: sequence composition determines dense attention work in each data-parallel microbatch, while token routing determines sparse expert work on expert-parallel ranks.
By Yibo Shen, Xudong Han, Xiaowei Zhu, Gen Li, Zhenxuan Pan
arXiv:2606. 01281v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs).
By Yixiu Mao, Yun Qu, Qi Wang, Heming Zou, Xiangyang Ji
arXiv:2609.23085v1 Announce Type: cross
Abstract: Large language models (LLMs) and agentic AI systems are creating rapidly growing inference energy demands as model sizes grow and reasoning trajector...
By Muhammad Abdur Rab Siddiqui, Daniela Rojas, Chen Yang, Wenqi Cui, Yuanyuan Shi, Yize Chen
CounterRoute is an online reinforcement‑learning framework that jointly learns how to route a language model’s reasoning between a ‘think’ and a ‘direct answer’ mode, using a single shared policy derived from a dual‑mode checkpoint. It employs counterfactual rollouts to credit routing decisions and a curriculum that starts with forced dual‑mode rollouts before shifting to self‑routed updates, achieving better accuracy‑efficiency trade‑offs across nine benchmarks. The method reduces generated tokens by up to 51% on Qwen3‑8B while improving macro‑average accuracy, and its routing strategy generalizes to unseen coding, science, and commonsense tasks.
By Ruochen Jiao, Besnik Fetahu, Zhenyu Shi, Priyanka Nigam
The paper introduces a history‑aware offline reinforcement learning policy that predicts iterative cost weights for routing in dense integrated circuit designs. By incorporating a lightweight LSTM and additional router features, the policy retains sequence context and improves convergence across various placement densities and guide qualities. Integrated into any cost‑based router with minimal changes, the approach reduces design rule violations by an average of 92% and cuts runtime by 10%.
arXiv:2607. 08780v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models activate only a sparse subset of experts per token, yet consecutive tokens frequently activate different experts -- causing constant weight swapping between slow storage and fast memory on edge devices.
By Ali Kayyam
arXiv:2607. 22465v1 Announce Type: cross Abstract: Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI.
By Ritik Raj, Souvik Kundu, Sarbartha Banerjee, Dheemanth Joshi, Ishita Vohra, Tushar Krishna
arXiv:2606. 29758v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) for Large Language Models increasingly relies on critic-free methods as a practical alternative to actor--critic training.
By Doo Hwan Hwang, Kee-Eung Kim