arXiv Machine Learning

Self-Routed Tensor Adapters for Parameter-Efficient Universal Visual Adaptation

arXiv:2608. 16384v1 Announce Type: cross Abstract: Universal visual representations require adaptation mechanisms that adapt across heterogeneous domains without fragmenting knowledge into domain-specific modules.

arXiv Computer Vision
2d ago

Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation

The paper investigates whether the sparsity of Mixture-of-Experts (MoE) models leads to intrinsic semantic organization across modalities and domains. It shows that experts naturally specialize semantically even without explicit modular training. The authors propose ExpertLens, a data‑free method that decodes router weights to identify domain‑specialized experts, enabling selective fine‑tuning that matches or exceeds full fine‑tuning while updating only 21.7–47.0% of parameters and achieving a 4.0× speedup, outperforming LoRA in both performance and efficiency.

By Damiano Marsili, Raphi Kang, Aditya Mehta, Pietro Perona, Georgia Gkioxari
arXiv Computation and Language
Aug 31

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

PRISM is a training‑free framework that efficiently selects visual instruction data for multimodal large language models by addressing the anisotropy in visual feature distributions, which causes a Global Semantic Drift. By implicitly re‑centering visual semantics, PRISM removes the influence of global background features, cutting data‑selection and model‑tuning time to 30% of conventional pipelines while improving performance across eight multimodal and three language benchmarks, achieving a 101.7% relative gain over baseline models.

By Jinhe Bi, Aniri, Zengjie Jin, Yifan Wang, Danqi Yan, Wenke Huang, Xiaowen Ma, Sikuan Yan, Artur Hecker, Mang Ye, Xun Xiao, Hinrich Schuetze, Volker Tresp, Yunpu Ma
Hugging Face Trending Papers
Aug 4

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solution, we identify a fundamental bottleneck in current methods: the reliance on static rank configurations.

arXiv Computer Vision
Aug 31

Task-State Adaptation with Prototype Memory for Multi-Task Dense Prediction

The paper introduces MemMTL, a multi‑task dense prediction framework that uses a compact task state derived from global visual context and refines it via a learnable prototype memory. This refined state informs task‑conditioned expert logits, which are combined with token‑level logits and routed through a sparse top‑k selection over a shared local expert bank. A task‑agnostic residual bank offers a common adaptation path, and both paths are added to the backbone feature before task‑specific prediction. The authors outline an evaluation protocol on NYUD‑v2 and PASCAL‑Context using SAM 3 and ViT‑L backbones to assess predictive quality, computational cost, and the contributions of task‑state conditioning, prototype retrieval, and sparse routing.

By Yangyang Xu, Haobo Yuan, Yuzhu Wang, Duo Su, Xi Ye, Yibo Yang, Jun Zhu