Hugging Face Trending Papers

Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing

The paper introduces a two‑level internal readout for mixture‑of‑experts reasoning models. First, it compresses the model’s reasoning states into a 64‑axis semantic frame (J64) that reveals process states not captured by the emitted trace, improving held‑out AUC by 0.096–0.135. Second, it reconstructs this frame from native expert‑routing statistics (R64), achieving high correlation with J64 and preserving most of its predictive gain while enabling low‑overhead, test‑time decision making and improved routing policies.

arXiv AI
Aug 19

Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing

The paper introduces a two‑level readout for mixture‑of‑experts reasoning models. First, it compresses the model’s internal reasoning states into a 64‑dimensional semantic frame (J64) that reveals process dynamics beyond the emitted trace. Second, it reconstructs this frame from native expert‑routing statistics (R64), achieving high correlation and preserving most predictive gains while enabling low‑overhead, test‑time decision making.

By Kang Chen, Sihan Zhao, Yixin Cao, Yugang Jiang
arXiv AI
Sep 7

When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models

The paper examines how expert pruning—removing low‑importance experts in Mixture‑of‑Experts models—fails when the router is over‑dispersed, a condition caused by aggressive load‑balancing that spreads tokens nearly uniformly across experts. In this regime, traditional importance signals from router probabilities collapse, making perplexity an unreliable predictor of downstream accuracy; for example, the lowest‑perplexity pruning on gpt‑oss‑20B harms mathematical reasoning while the highest‑perplexity pruning preserves it. To address this, the authors introduce Minimax Expert Score Allocation (MESA), a domain‑aware method that iteratively boosts scores for the most affected domain, achieving minimal worst‑case degradation across domains and outperforming baseline pruning strategies on multiple benchmarks while reducing memory usage.

By Berkcan Kapusuzoglu, Connor Pryor, Sangwoo Cho, Supriyo Chakraborty, Shi-Xiong Zhang, Sambit Sahu, Milind Naphade
Hugging Face Trending Papers
Jul 27

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU.

arXiv Machine Learning
Aug 27

Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing

The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.

By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng
arXiv Machine Learning
Sep 22

Efficient Mixture-of-Experts with Speculative Decoding via Expert Coactivation

The paper studies how the design of Mixture-of-Experts (MoE) routers affects inference speed when combined with Speculative Decoding (SD). It shows that routers promoting high expert coactivation reduce memory transfer costs and improve runtime. By integrating a global load‑balancing loss, shared experts, a consistency loss, and an autoregressive expert selection mechanism, the authors achieve a 21% throughput gain over baseline MoEs while preserving accuracy.

By Kumari Nishu, Han-Byul Kim, Santosh Chilkunda, Maxwell Horton, Arnav Kundu, Mohammad Samragh, Lauren Hannah, Mohammad Sekhavat, Nikhil Bhendawade, Manuel Ciosici, Iman Mirzadeh, Keivan Alizadeh Vahid, David Harrison, Irina Belousova, Mehrdad Farajtabar, Minsik Cho