Hugging Face Trending Papers

Routing in Gradient Space: Balanced Usage Is Not Expert Specialization

arXiv Machine Learning
Jul 15

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs

arXiv:2602. 19938v2 Announce Type: replace Abstract: Sparse Mixture-of-Experts (SMoE) architectures are increasingly used to scale large language models efficiently, delivering strong accuracy under fixed compute budgets.

By Zijie Liu, Jie Peng, Jinhao Duan, Zirui Liu, Kaixiong Zhou, Mingfu Liang, Luke Simon, Xi Liu, Zhaozhuo Xu, Tianlong Chen
arXiv AI
Sep 7

When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models

The paper examines how expert pruning—removing low‑importance experts in Mixture‑of‑Experts models—fails when the router is over‑dispersed, a condition caused by aggressive load‑balancing that spreads tokens nearly uniformly across experts. In this regime, traditional importance signals from router probabilities collapse, making perplexity an unreliable predictor of downstream accuracy; for example, the lowest‑perplexity pruning on gpt‑oss‑20B harms mathematical reasoning while the highest‑perplexity pruning preserves it. To address this, the authors introduce Minimax Expert Score Allocation (MESA), a domain‑aware method that iteratively boosts scores for the most affected domain, achieving minimal worst‑case degradation across domains and outperforming baseline pruning strategies on multiple benchmarks while reducing memory usage.

By Berkcan Kapusuzoglu, Connor Pryor, Sangwoo Cho, Supriyo Chakraborty, Shi-Xiong Zhang, Sambit Sahu, Milind Naphade
arXiv Computation and Language
Aug 28

Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs

MetaNet is a support‑set controller that predicts, for each layer of a Mixture‑of‑Experts model, an expert‑retention threshold and a bounded routing bias while keeping the backbone, experts, and router frozen. On DeepSeek‑MoE‑16B‑Chat, MetaNet offers a tunable trade‑off between accuracy and expert activation: a conservative setting activates 3.61 experts on average (40% fewer than a fixed k=6) with comparable MMLU accuracy, whereas an aggressive setting activates only 2.28 experts (62% fewer) with a modest accuracy drop. The MMLU‑trained controller also transfers to C‑Eval, activating 2.90 experts on average (52% fewer than fixed k=6) at 0.386 accuracy.

By Rongfeng Wang, Shichao Weng, Zhiqiang Wang, Xinyu Liu, Yang Yi, Peilong Zhou, Hongwei Tang
arXiv AI
4d ago

Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

The paper introduces two token-error supervision methods for sparse mixture-of-experts (MoE) large language models, aligning routing affinities with token-level cross-entropy loss. The first method predicts an error score per expert and uses it to adjust affinities before top‑K selection, while the second directly aligns router affinities to the model’s objective without extra heads or inference changes. Experiments on Granite and ARC‑Challenge show accuracy gains of about 2.3–2.94 percentage points over a parameter‑matched baseline, preserving the native sparse execution budget.

By Yury Nahshan, Nati Daniel, Jacob Goldberger, Yoli Shavit
arXiv Machine Learning
1d ago

From Task Mixtures to Specialized Experts

The paper investigates federated learning where each client’s data consists of unknown mixtures of distinct tasks, a scenario termed compound heterogeneity. It shows that when tasks share a common feature geometry, the optimal model for a mixed client is a convex combination of task‑specific models, motivating input‑dependent routing to specialized experts. The authors propose FedSEE, a method that recovers task experts via a convex program and achieves better performance than baselines, reducing negative transfer by 2.9 points overall and 3.7 points for the worst‑served quartile.

By Hojat Allah Salehi, Mehrdad Mahdavi, Andrew Arash Mahyari, M. Hadi Amini
Hugging Face Trending Papers
Aug 10

MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts

Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly rely on coarse prompt-level differentiation without parameter adaptation for diverse subtasks, resulting in insufficient inter-agent heterogeneity and limited specialized capability that bottleneck performance on tasks with complex requirements.