arXiv Machine Learning

$\alpha$Transfer: Coefficient Transfer for Efficient Model Merging

arXiv AI
Sep 18

Riemannian--Lorentz Fusion of Vision Transformers and State-Space Models

The paper introduces Riemannian–Lorentz Parameter Fusion (RLPF), a method for merging a Vision Transformer and a state‑space model without gradient descent. RLPF aligns parameter groups by semantic role, projects them onto a common coordinate system, lifts selected coordinates to the Lorentz hyperboloid, computes a regularized geodesic barycenter, and decodes the result back into the two branches, with a learned gate combining their logits. The resulting fine‑tuned system achieves 82.37 % on CIFAR‑10, 75.04 % on Oxford‑IIIT Pet, and 78.58 % top‑1 accuracy on ImageNet‑1K, surpassing the best‑parent accuracies of 76.54 %, 71.42 %, and 76.42 % respectively.

By Badri N. Patro, Vijay S. Agneeswaran
arXiv Machine Learning
Sep 22

Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion

Merge++ is a post‑hoc refinement technique for model merging that synthesizes task‑representative images by inverting expert checkpoints and then distills expert knowledge into a merged model. It operates without any additional data beyond the checkpoints and can be applied universally across existing weight‑space merging algorithms. Experiments show consistent improvements, with average gains of +2 to +8 points and up to +25.9 on specific configurations.

By Aditya Pola, Vineeth N. Balasubramanian
arXiv Machine Learning
Jun 19

Model soups need only one ingredient

arXiv:2602. 09689v2 Announce Type: replace Abstract: Fine-tuning large pre-trained models on a target distribution often improves in-distribution (ID) accuracy, but at the cost of out-of-distribution (OOD) robustness as representations specialize to the fine-tuning data.

By Alireza Abdollahpoorrostam, Nikolaos Dimitriadis, Adam Hazimeh, Pascal Frossard
arXiv AI
Sep 4

One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging

The paper investigates weight‑space merging of independently fine‑tuned multilingual machine translation models. Experiments show that merging is more successful when models share a target language, yet it still cannot match the peak performance of language‑specific checkpoints. When target languages differ, performance drops sharply, and analysis reveals that overlapping neuron activation and incompatible upper‑layer geometries cause these failures.

By Baban Gain, Trilok Nath Singh, Asif Ekbal
arXiv Computation and Language
Aug 31

Pruning Laws for Large Language Models

arXiv:2504.04342v2 Announce Type: replace Abstract: Scaling up model parameters and training data consistently improves the performance of large language models (LLMs), but at the cost of rapidly gro...

By Ayan Sengupta, Siddhant Chaudhary, Tanmoy Chakraborty
arXiv AI
6d ago

Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models

Architectural Sampling is a training‑free technique that improves test‑time scaling for frozen vision‑language models by generating diverse candidates through distinct forward computations. It achieves this by reusing selected blocks of decoder layers, varying block location and repetition count, thereby creating computational diversity without updating weights or adding parameters. Experiments on five Qwen checkpoints across twelve multimodal benchmarks show that this method raises pass@9 by an average of 6.58 percentage points over standard temperature sampling, with early‑layer reuse delivering the strongest gains and lower lexical overlap in the generated candidates.

By Akshit Singh, Shyam Marjit, Wei Lin, Leonid Karlinsky, M. Jehanzeb Mirza