arXiv AI

Signature-Guided Capacity Occupancy for Dense Expert Merging

arXiv:2608. 09201v1 Announce Type: new Abstract: Dense expert merging combines domain-specialized language models into one single checkpoint, typically by admitting task-vector support in weight space.

arXiv Machine Learning
Sep 22

Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion

Merge++ is a post‑hoc refinement technique for model merging that synthesizes task‑representative images by inverting expert checkpoints and then distills expert knowledge into a merged model. It operates without any additional data beyond the checkpoints and can be applied universally across existing weight‑space merging algorithms. Experiments show consistent improvements, with average gains of +2 to +8 points and up to +25.9 on specific configurations.

By Aditya Pola, Vineeth N. Balasubramanian
arXiv Machine Learning
Sep 21

IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

IntBMoE introduces a block‑conditioned mixture‑of‑experts that decouples participation, execution, and materialization by combining dense expert composition with sparse block execution. Each internal layer uses a lightweight hypernetwork to merge all expert bases into a single composed expert, while a router selects only a few blocks per token, keeping compute and memory costs low. Experiments on image classification, language modeling, and sequential recommendation demonstrate consistent performance gains, and the model is deployed in AMap’s generative recommendation system, improving UVCTR by 2.4% in online A/B tests.

By Ran Cheng, Longfei Xu, Zheng Liu, Kaikui Liu, Xiangxiang Chu
arXiv Machine Learning
Sep 22

CAMFT: Conflict-Aware Mergeable Fine-Tuning for Large Language Models

CAMFT is a Conflict‑Aware Mergeable Fine‑Tuning method designed to make task adaptation efficient and merge‑aware for large language models. Unlike existing approaches that only resolve parameter conflicts after fine‑tuning, CAMFT shapes mergeability during training by guiding each task to update sparse coordinates with lower cross‑task conflict. Experiments show that CAMFT outperforms standard fine‑tuning baselines in multi‑task merging scenarios.

By Jingang Zhou, Haiyang Guo, Yuan Ma, Han Zhu, Xu-Yao Zhang
arXiv AI
Jun 2

DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts

arXiv:2606. 01062v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have become a leading approach for decoupling parameter count from computational cost in large language models, yet effectively scaling MoE performance remains a challenge.

By Jiarui Feng, Hanqing Zeng, Karish Grover, Ruizhong Qiu, Yinglong Xia, Qiang Zhang, Qifan Wang, Ren Chen, Dongqi Fu, Jiayi Liu, Zhoukai Zhao, Xiangjun Fan, Benyu Zhang, Yixin Chen
arXiv Computation and Language
Aug 28

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

The paper investigates three fusion paradigms—Merge, Mix RL, and multi‑teacher on‑policy distillation (MOPD)—for consolidating reinforcement learning with verifiable rewards (RLVR) across multiple domains. Experiments across model scales and a multi‑domain benchmark show that while overall performance differences are small, significant gaps can appear on specific tasks, and each method exhibits distinct training dynamics and constraints. Practical guidelines are offered: Merge for cheap fusion when experts exist, Mix RL for unified training with adjustable domain mixtures, and MOPD when preserving domain‑specific gains is paramount.

By Siye Wu, Kai Yang, Yuchen Cai, Xin Xu, Peng-Yuan Wang, Jiaxuan Wang, Jiashun Liu, Jiafei Lyu, Yangkun Chen, Saiyong Yang, Yanghua Xiao
arXiv AI
Aug 11

UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models

arXiv:2608. 08627v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still stores and routes over its full expert bank.

By Lei Xin, Bin Gu, Peize Li, Zitong Wang, Jianbo Zhao, Changjiang Jiang, Yanyue Xie, Chao Huang, Xuyang Zhao, Zunhai Su, Fanhu Zeng, Zhenglun Kong