arXiv:2607. 26052v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) variants of Low-Rank Adaptation (LoRA) route every token to a fixed number of experts $k$.
By Tom Saliencro, Rohan Desai, Priya Nair, Maya Lindqvist, Daniel Whitmore
The paper introduces READ, a method for composing low‑rank adapters (LoRA) in large language models. By rewriting each adapter into a balanced canonical form and enforcing a one‑directional coupling, READ allows new skills to read but never write into the output subspaces of existing skills, eliminating interference. Experiments on four benchmark suites and two model families show that READ consistently outperforms existing baselines, improving SuperGLUE scores by over twenty points and domain suite scores by more than seven points.
By Zeyan Li, Panqi Yang, Qirong Guo, Shengda Zhuo, SIyuan Qiu, Hu Xu, Chun Li, Jianfeng Xu
arXiv:2608. 07814v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many expert weight matrices.
By Inesh Chakrabarti, Sourjya Roy, Bowen Bao, Thiago Crepaldi, Spandan Tiwari, Ashish Sirasao
arXiv:2608. 07890v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models decouple total parameters from per-token compute, but deployment still requires storing every expert.
By Ali Janati, Kaoutar El Maghraoui, Xinyi Luo, Wenyuan Shen, Owen Zou, Yankai Mao
arXiv:2608. 08853v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) routers commonly use the same scores both to select experts and to weight their already-computed outputs.
By Zongfei Li
arXiv:2606. 31413v1 Announce Type: new Abstract: Composing independently trained LoRA adapters into a single large language model is useful for multi-domain adaptation, especially when the original training data cannot be shared.
By Seyed Alireza Molavi, Zhan Su, Yan Hu, Peyman Sheikholharam Mashhadi, Stefan Byttner, Prayag Tiwari
IntBMoE introduces a block‑conditioned mixture‑of‑experts that decouples participation, execution, and materialization by combining dense expert composition with sparse block execution. Each internal layer uses a lightweight hypernetwork to merge all expert bases into a single composed expert, while a router selects only a few blocks per token, keeping compute and memory costs low. Experiments on image classification, language modeling, and sequential recommendation demonstrate consistent performance gains, and the model is deployed in AMap’s generative recommendation system, improving UVCTR by 2.4% in online A/B tests.
By Ran Cheng, Longfei Xu, Zheng Liu, Kaikui Liu, Xiangxiang Chu
JevSoup is a training‑free framework that separates System One expert routing from System Two execution for low‑rank adaptation (LoRA) models. It selects two experts based solely on input and expert descriptions, retains the leading expert’s update, projects the second onto the orthogonal complement of the first update’s row space, and combines them with equal weights. On 14 PorTAL tasks and three Qwen3 scales, JevSoup improves task‑macro accuracy by up to 1.19 % and sample‑micro accuracy by up to 1.21 % over the strongest evaluated baselines.
By Xiuying Wang, Jiahua Cheng, Shuotian Li, Yufan Cheng, Bowen Deng, Zhexuan Bai, Yichen Li
LoRA-TSD introduces a new optimizer for low‑rank adaptation (LoRA) that treats each update as a tangent vector on the fixed‑rank matrix manifold and applies a Muon‑style spectral‑norm steepest‑descent step within that tangent space. The method avoids costly full‑matrix operations and offers a retraction that is up to 2.8× cheaper than previous manifold approaches. The authors prove that their surrogate recovers LoRA‑Pro, identify the Riemannian gradient as the natural stationarity measure, and provide the first global convergence guarantees for both LoRA‑Pro and LoRA‑TSD, achieving superior performance across multiple benchmarks with Llama and Qwen models.
By Dmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov
ACE introduces a training‑free, calibration‑free framework for adaptive expert skipping in Mixture‑of‑Experts LLMs. It combines a Global Spectral Proxy that estimates global transformation capacity with a Router‑Conditioned Refinement that builds expert‑specific direction prototypes, enabling the model to skip low‑contribution experts while always keeping the top‑1 expert. Offline computation of expert statistics leaves only lightweight table lookups during inference, and experiments on three MoE‑based LLMs show ACE outperforms static and dynamic baselines, especially at high skipping ratios.
By Zukang Xu, Zhixiong Zhao, Xing Hu, Jiangyong Yu, Houji Wen, Jun Li, Zhe Jiang, Dawei Yang
arXiv:2601. 13020v2 Announce Type: replace-cross Abstract: Continual instruction tuning (CIT) requires multimodal large language models (MLLMs) to adapt to a stream of tasks without forgetting prior capabilities.
By Zhiyan Hou, Haiyun Guo, Haokai Ma, Yandu Sun, Yonghui Yang, Jinqiao Wang
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that treats every LoRA step as a tangent vector of the fixed-rank matrix manifold and takes the spectral-norm steepest-descent step of Muon inside that tangent space, mapping the result back to the factors through a retraction native to the LoRA parametrization.