arXiv AI
1d ago

VALSE: Vertical Adaptive Layer Skipping for Efficient Inference in Large Language Models

The paper introduces VALSE, a vertical adaptive layer skipping technique for large language models. It provides a theoretical framework proving an Expected FLOPs formula, showing that skip-layer models are a strict subset yet meaningful approximation of full-layer models, and revealing a duality between VALSE and Mixture-of-Experts architectures. VALSE uses a lightweight difficulty estimator to selectively skip redundant layers per input, enabling efficient inference by activating only necessary depth.

By Jia-Dong Zhang
arXiv AI
Jun 4

L$^3$: Large Lookup Layers

arXiv:2601. 21461v3 Announce Type: replace-cross Abstract: Modern sparse language models typically achieve sparsity through Mixture-of-Experts (MoE) layers, which dynamically route tokens to dense MLP "experts.

By Albert Tseng, Christopher De Sa