arXiv AI

Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs

arXiv:2606. 31413v1 Announce Type: new Abstract: Composing independently trained LoRA adapters into a single large language model is useful for multi-domain adaptation, especially when the original training data cannot be shared.

arXiv Machine Learning
Jun 5

IR3DE: A Linear Router for Large Language Models

arXiv:2606. 06098v1 Announce Type: cross Abstract: Foundational Large Language Models (LLMs) demonstrate proficiency on a wide range of general tasks, and achieve remarkable results on various specialized tasks via domain-expert LLMs.

By Eros Fan\`i, O\u{g}uzhan Ersoy
arXiv Machine Learning
Aug 27

Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing

The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.

By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng
arXiv AI
6d ago

JevSoup: System-One Routing for Training-Free LoRA Composition

JevSoup is a training‑free framework that separates System One expert routing from System Two execution for low‑rank adaptation (LoRA) models. It selects two experts based solely on input and expert descriptions, retains the leading expert’s update, projects the second onto the orthogonal complement of the first update’s row space, and combines them with equal weights. On 14 PorTAL tasks and three Qwen3 scales, JevSoup improves task‑macro accuracy by up to 1.19 % and sample‑micro accuracy by up to 1.21 % over the strongest evaluated baselines.

By Xiuying Wang, Jiahua Cheng, Shuotian Li, Yufan Cheng, Bowen Deng, Zhexuan Bai, Yichen Li
Hugging Face Trending Papers
Aug 4

MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation

Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capacity by storing multiple full LoRA experts, causing adapter storage to grow linearly with the number of experts and restricting adaptation to a fixed expert pool. We ask whether MoE-based PEFT can produce instance-specific adaptations without explicitly storing a separate LoRA module for each expert.

arXiv AI
Jun 2

DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts

arXiv:2606. 01062v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have become a leading approach for decoupling parameter count from computational cost in large language models, yet effectively scaling MoE performance remains a challenge.

By Jiarui Feng, Hanqing Zeng, Karish Grover, Ruizhong Qiu, Yinglong Xia, Qiang Zhang, Qifan Wang, Ren Chen, Dongqi Fu, Jiayi Liu, Zhoukai Zhao, Xiangjun Fan, Benyu Zhang, Yixin Chen
arXiv Computation and Language
4d ago

MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation

arXiv:2608.03275v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve...

By Yiming Zeng, Lei Lu, Zexin Li, Zhuochun Li, Dehai Min, Shuoqiu Li, Shuyi Liao, Xidong Wu, Zeyu Zhang, Minmei Wang, Yu Zhao, Tingting Yu, Shangqian Gao