arXiv:2608. 02989v1 Announce Type: new Abstract: Speculative decoding verifies a tree of draft tokens in one target-model forward pass.
By Shuang Liang (Mark), Hao (Mark), Chen, Zhiwen Mo, Qianzhou Wang, Guoyu Li, Lingxiao Ma, Wayne Luk
ACE introduces a training‑free, calibration‑free framework for adaptive expert skipping in Mixture‑of‑Experts LLMs. It combines a Global Spectral Proxy that estimates global transformation capacity with a Router‑Conditioned Refinement that builds expert‑specific direction prototypes, enabling the model to skip low‑contribution experts while always keeping the top‑1 expert. Offline computation of expert statistics leaves only lightweight table lookups during inference, and experiments on three MoE‑based LLMs show ACE outperforms static and dynamic baselines, especially at high skipping ratios.
By Zukang Xu, Zhixiong Zhao, Xing Hu, Jiangyong Yu, Houji Wen, Jun Li, Zhe Jiang, Dawei Yang
arXiv:2608. 02528v1 Announce Type: new Abstract: Mixtures of low-rank adaptation experts increase parameter-efficient capacity by routing each input through a subset of adapters.
By Tom Saliencro, Rohan Desai, Priya Nair, Maya Lindqvist, Daniel Whitmore
The paper introduces VANE, a new routing strategy for Mixture-of-LoRA-experts that scores experts based on the expected loss reduction of their updates rather than token similarity. By decomposing each expert into a reader and writer, VANE evaluates all reader–writer pairs using a low‑rank compass that predicts the descent direction, enabling efficient top‑k selection with additive gates. Experiments on Llama‑3 models show VANE outperforms twelve PEFT and MoE‑LoRA baselines while using fewer trainable parameters and providing router scores that better reflect expert usefulness.
By Priya Nair, Lukas Brenner, Maya Lindqvist, Daniel Whitmore, Wen-Hsuan Liu, Tom Saliencro, Amara Okonkwo, Rohan Desai
arXiv:2608. 07890v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models decouple total parameters from per-token compute, but deployment still requires storing every expert.
By Ali Janati, Kaoutar El Maghraoui, Xinyi Luo, Wenyuan Shen, Owen Zou, Yankai Mao
MetaNet is a support‑set controller that predicts, for each layer of a Mixture‑of‑Experts model, an expert‑retention threshold and a bounded routing bias while keeping the backbone, experts, and router frozen. On DeepSeek‑MoE‑16B‑Chat, MetaNet offers a tunable trade‑off between accuracy and expert activation: a conservative setting activates 3.61 experts on average (40% fewer than a fixed k=6) with comparable MMLU accuracy, whereas an aggressive setting activates only 2.28 experts (62% fewer) with a modest accuracy drop. The MMLU‑trained controller also transfers to C‑Eval, activating 2.90 experts on average (52% fewer than fixed k=6) at 0.386 accuracy.
By Rongfeng Wang, Shichao Weng, Zhiqiang Wang, Xinyu Liu, Yang Yi, Peilong Zhou, Hongwei Tang