arXiv AI

Patch Rebirth: Fast and Transferable Model Inversion of Vision Transformers

Hugging Face Trending Papers
Aug 11

Grid-Preserving Knowledge Distillation: Transferring Convolutional Inductive Bias to Vision Transformers under Data Scarcity

Vision Transformers underperform convolutional networks when training data is scarce, and distilling convolutional inductive biases from a CNN teacher is an effective remedy that leaves the deployed model unchanged. General-purpose feature distillation, however, transfers little in this setting.

arXiv Computation and Language
Aug 31

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

PRISM is a training‑free framework that efficiently selects visual instruction data for multimodal large language models by addressing the anisotropy in visual feature distributions, which causes a Global Semantic Drift. By implicitly re‑centering visual semantics, PRISM removes the influence of global background features, cutting data‑selection and model‑tuning time to 30% of conventional pipelines while improving performance across eight multimodal and three language benchmarks, achieving a 101.7% relative gain over baseline models.

By Jinhe Bi, Aniri, Zengjie Jin, Yifan Wang, Danqi Yan, Wenke Huang, Xiaowen Ma, Sikuan Yan, Artur Hecker, Mang Ye, Xun Xiao, Hinrich Schuetze, Volker Tresp, Yunpu Ma