Hugging Face Trending Papers

VALSE: Vertical Adaptive Layer Skipping for Efficient Inference in Large Language Models

arXiv AI
1d ago

VALSE: Vertical Adaptive Layer Skipping for Efficient Inference in Large Language Models

The paper introduces VALSE, a vertical adaptive layer skipping technique for large language models. It provides a theoretical framework proving an Expected FLOPs formula, showing that skip-layer models are a strict subset yet meaningful approximation of full-layer models, and revealing a duality between VALSE and Mixture-of-Experts architectures. VALSE uses a lightweight difficulty estimator to selectively skip redundant layers per input, enabling efficient inference by activating only necessary depth.

By Jia-Dong Zhang
arXiv AI
Jun 4

L$^3$: Large Lookup Layers

arXiv:2601. 21461v3 Announce Type: replace-cross Abstract: Modern sparse language models typically achieve sparsity through Mixture-of-Experts (MoE) layers, which dynamically route tokens to dense MLP "experts.

By Albert Tseng, Christopher De Sa
arXiv AI
Jun 3

DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training

arXiv:2512. 13996v3 Announce Type: replace Abstract: Sparse Mixture-of-Experts architectures are essential for scaling model capacity efficiently, yet the standard Top-$k$ routing imposes a rigid sparsity pattern that ignores the intrinsic variance in token difficulty and layer-specific computational needs.

By Can Jin, Hongwu Peng, Mingcan Xiang, Qixin Zhang, Xiangchi Yuan, Amit Hasan, Ohi Dibua, Yifan Gong, Yan Kang, Dimitris N. Metaxas
arXiv AI
Sep 17

MoRE: Mixture of Reused Experts

MoRE: Mixture of Reused Experts is a hybrid architecture that combines Mixture-of-Experts (MoE) with weight‑sharing techniques. It shares expert pools across adjacent layers while each layer keeps its own router, and introduces lightweight depth embeddings to help shared experts differentiate layer contexts. Experiments on models ranging from 114 M to 1.15 B parameters show MoRE achieves lower perplexity and better downstream performance than standard MoEs and other weight‑sharing models, with only minimal changes to existing MoE implementations.

By Eric S. Qiu, Utku Umur Acikalin, Justin Lovelace, Christian Belardi, Arjun B. Mulchandani, Carla P. Gomes, Kilian Q. Weinberger
arXiv Machine Learning
Sep 21

IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

IntBMoE introduces a block‑conditioned mixture‑of‑experts that decouples participation, execution, and materialization by combining dense expert composition with sparse block execution. Each internal layer uses a lightweight hypernetwork to merge all expert bases into a single composed expert, while a router selects only a few blocks per token, keeping compute and memory costs low. Experiments on image classification, language modeling, and sequential recommendation demonstrate consistent performance gains, and the model is deployed in AMap’s generative recommendation system, improving UVCTR by 2.4% in online A/B tests.

By Ran Cheng, Longfei Xu, Zheng Liu, Kaikui Liu, Xiangxiang Chu