arXiv Machine Learning

Training-free Task Classification for Multi-Task Model Merging

arXiv:2606. 22589v2 Announce Type: replace Abstract: Ever since the advent of foundation models and the pre-training-finetuning paradigm, there have been numerous efforts to merge multiple task-specific experts into a single multi-task model.

Hugging Face Trending Papers
Jun 25

Learning to Recover Task Experts from a Multi-Task Merged Model

Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly storage and loading of redundant expert components at inference.

arXiv Machine Learning
Aug 27

Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement

The paper introduces a new multi‑task model‑merging framework that tackles task interference by projecting task vectors into a high‑dimensional sparse feature space using Sparse Autoencoders, enabling feature‑level disentanglement before fusion. It also proposes a lightweight Group‑Ranked Zeroth‑Order Optimizer to identify task‑critical layers for selective merging, reducing computational overhead. Experiments on Qwen2.5‑1.5B and Qwen2.5‑7B show consistent performance gains over several baselines across reasoning, code generation, instruction following, and general knowledge tasks, with a 2.78% improvement in a highly conflicting four‑task setting.

By Yihang Zhang, Shengke Sun, Junjie Wen, Feng Zeng
arXiv AI
Sep 15

Task-Aware Federated Fine-Tuning for MoE-based Large Language Models

The paper introduces FedTAR, a task-aware federated fine‑tuning approach for Mixture‑of‑Experts (MoE) large language models. FedTAR links local client updates to task preferences using routing outputs and Singular Value Decomposition to extract low‑dimensional task coordinates and update directions. It then aggregates updates within and across task clusters, reconstructing the final update to preserve expert specialization and reduce interference, achieving state‑of‑the‑art performance on four benchmark tasks under non‑IID settings.

By Tingqi Wang, Hongyu Ke, Haoxin Wang, Rafal Angryk, Zhipeng Cai
arXiv AI
Aug 25

From Isolation to Alignment: Unified LoRA for Efficient Multi-Task Learning

The paper introduces Align‑LoRA, a unified LoRA framework for multi‑task learning that replaces complex, isolated adapter designs with a single‑adapter model enhanced by a higher rank and an explicit alignment loss. It demonstrates that a router‑free, multi‑head model with high inter‑head redundancy can outperform more elaborate baselines, and that a unified LoRA can achieve competitive performance while enabling weight merging and zero inference latency. Extensive experiments and theoretical analysis confirm that Align‑LoRA surpasses prevailing approaches, offering a simpler, production‑friendly paradigm for parameter‑efficient fine‑tuning of large language models.

By Jinda Liu, Yi Chang, Yuan Wu
arXiv Computer Vision
Sep 11

Task Alignment: A Simple Proxy for Practical Model Merging Across Diverse Vision Tasks

The paper introduces the task alignment proxy, a method that accelerates hyperparameter selection for merging models fine‑tuned on diverse vision tasks. It addresses the challenge of training heterogeneous decoders, which makes traditional downstream performance evaluation costly. By using the proxy, the authors demonstrate that model merging can be applied efficiently to multi‑task vision models beyond CLIP‑based classification.

By Pau de Jorge, C\'esar Roberto de Souza, Bj\"orn Michele, Mert B\"ulent Sar{\i}y{\i}ld{\i}z, Philippe Weinzaepfel, Florent Perronnin, Diane Larlus, Yannis Kalantidis