arXiv AI

Lexicographic Multi-Objective On-Policy Distillation

arXiv AI
Jun 2

SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting

arXiv:2604. 10688v2 Announce Type: replace-cross Abstract: On-policy reinforcement learning has become the dominant paradigm for reasoning alignment in large language models, yet its sparse, outcome-level rewards make token-level credit assignment notoriously difficult.

By Binbin Zheng, Xing Ma, Yiheng Liang, Jingqing Ruan, Xiaoliang Fu, Kepeng Lin, Benchang Zhu, Ke Zeng, Xunliang Cai
arXiv Computation and Language
Aug 28

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

The paper investigates three fusion paradigms—Merge, Mix RL, and multi‑teacher on‑policy distillation (MOPD)—for consolidating reinforcement learning with verifiable rewards (RLVR) across multiple domains. Experiments across model scales and a multi‑domain benchmark show that while overall performance differences are small, significant gaps can appear on specific tasks, and each method exhibits distinct training dynamics and constraints. Practical guidelines are offered: Merge for cheap fusion when experts exist, Mix RL for unified training with adjustable domain mixtures, and MOPD when preserving domain‑specific gains is paramount.

By Siye Wu, Kai Yang, Yuchen Cai, Xin Xu, Peng-Yuan Wang, Jiaxuan Wang, Jiashun Liu, Jiafei Lyu, Yangkun Chen, Saiyong Yang, Yanghua Xiao
arXiv AI
Sep 17

Higher-order pruning of experts in mixture-of-experts language models

The paper introduces HOPE, a second‑order pruning method for Mixture‑of‑Experts language models that accounts for cooperative interactions between experts. Unlike first‑order methods such as REAP, HOPE derives an objective that provably bounds pruning error and is shown to outperform baselines across three large MoE models, multiple calibration sets, and diverse benchmarks, especially at high pruning rates and on agentic tasks. The results demonstrate that preserving expert interactions allows aggressive compression with minimal performance loss on complex workloads.

By Alex M. Tseng, Prannay Kaul, Luca Zancato, Wei Xia, Stefano Soatto