arXiv Machine Learning

PRIME: Mitigating Subgroup Optimization Competition in Shared CTR Top Networks with Plug-in Residual Input-Conditioned Mixture of Expert

PRIME is a plug‑in residual input‑conditioned mixture of experts that preserves the original dense prediction path while adding low‑rank, input‑dependent logit corrections. By initializing residuals to zero, PRIME matches the baseline dense model at training start and stabilizes conditional estimation with multi‑bag aggregation and EMA load biases. Experiments on Avazu and Criteo across 13 CTR architectures show modest AUC and LogLoss gains, with PRIME outperforming APG on FiBiNET and DCNv2 while using fewer parameters and lower latency.

arXiv Machine Learning
1d ago

The Conflict Between Logic and Memory: Learning Higher-Order Interactions in Shallow MLPs

The paper investigates how single‑hidden‑layer MLPs can fit training data yet fail to recover the underlying rule, focusing on higher‑order interactions and nuisance inputs. Using synthetic parity tasks, the authors benchmark different optimizers (SGD, Adam, Muon) and show that while all achieve perfect accuracy on second‑order interactions, performance drops sharply for higher orders, with Muon outperforming the others at fourth order. Experiments also reveal that freezing or removing nuisance‑related weights dramatically alters training outcomes, highlighting the role of nuisance learning in shaping the rules a shallow network can represent.

By Gongyue Zhang, Honghai Liu
arXiv AI
Jul 9

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation

arXiv:2607. 06601v1 Announce Type: cross Abstract: Conditional computation can decouple language model quality from per-token inference cost, yet leading techniques act on a single axis in isolation: Mixture-of-Experts (MoE) sparsifies the FFN, Mixture-of-Depths (MoD) skips whole transformer blocks, and KV-cache quantization compresses attention memory.

By Andrii Balashov, Olena Ponomarova
arXiv Computation and Language
Sep 1

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

arXiv:2608.30320v1 Announce Type: new Abstract: We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and a...

By Zihan Qiu, Zekun Wang, Xiao Li, Yanpeng Li, Yang Xu, Yixuan Wang, Huaqing Zhang, Rui Men, Bochao Mao, Chengruidong Zhang, Fan Zhou, Hao Luo, Haofeng Huang, Haoran Lian, Haoyan Huang, Hongqing Chen, Jianwei Zhang, Jing Xu, Junjie Wang, Langshi Chen, Liangyu Wang, Linlang Jiang, Man Yuan, Minmin Sun, Peng Jin, Siqi Zhang, Siyu Wang, Xingzhang Ren, Yakai Wang, Yi Zhang, Yiming Dong, Yizhong Cao, Yubo Ma, Yunfei Mao, Bo Zheng, Dayiheng Liu
arXiv AI
Jun 3

DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training

arXiv:2512. 13996v3 Announce Type: replace Abstract: Sparse Mixture-of-Experts architectures are essential for scaling model capacity efficiently, yet the standard Top-$k$ routing imposes a rigid sparsity pattern that ignores the intrinsic variance in token difficulty and layer-specific computational needs.

By Can Jin, Hongwu Peng, Mingcan Xiang, Qixin Zhang, Xiangchi Yuan, Amit Hasan, Ohi Dibua, Yifan Gong, Yan Kang, Dimitris N. Metaxas