arXiv AI By Xiuying Wang, Jiahua Cheng, Shuotian Li, Yufan Cheng, Bowen Deng, Zhexuan Bai, Yichen Li

JevSoup: System-One Routing for Training-Free LoRA Composition

Read the original on arXiv AI →

JevSoup is a training‑free framework that separates System One expert routing from System Two execution for low‑rank adaptation (LoRA) models. It selects two experts based solely on input and expert descriptions, retains the leading expert’s update, projects the second onto the orthogonal complement of the first update’s row space, and combines them with equal weights. On 14 PorTAL tasks and three Qwen3 scales, JevSoup improves task‑macro accuracy by up to 1.19 % and sample‑micro accuracy by up to 1.21 % over the strongest evaluated baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 15

Task-Aware Federated Fine-Tuning for MoE-based Large Language Models

The paper introduces FedTAR, a task-aware federated fine‑tuning approach for Mixture‑of‑Experts (MoE) large language models. FedTAR links local client updates to task preferences using routing outputs and Singular Value Decomposition to extract low‑dimensional task coordinates and update directions. It then aggregates updates within and across task clusters, reconstructing the final update to preserve expert specialization and reduce interference, achieving state‑of‑the‑art performance on four benchmark tasks under non‑IID settings.

By Tingqi Wang, Hongyu Ke, Haoxin Wang, Rafal Angryk, Zhipeng Cai
arXiv Machine Learning
Aug 27

Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing

The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.

By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng