Hugging Face Blog

🤗 PEFT welcomes new merging methods

Towards Data Science
Sep 28

Guided Merge Sort : An Optimized Sorting that Picks the Best from Ordinary and Multi-Way Merge Sort Algorithms

The article introduces Guided Merge Sort, an optimized sorting technique that combines elements of ordinary merge sort and multi‑way merge sort. It highlights how the use of the "goto" operator becomes essential in this approach. The post explains the algorithm’s design and its potential advantages over traditional methods.

By Tigran Hayrapetyan
arXiv Machine Learning
Sep 22

Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion

Merge++ is a post‑hoc refinement technique for model merging that synthesizes task‑representative images by inverting expert checkpoints and then distills expert knowledge into a merged model. It operates without any additional data beyond the checkpoints and can be applied universally across existing weight‑space merging algorithms. Experiments show consistent improvements, with average gains of +2 to +8 points and up to +25.9 on specific configurations.

By Aditya Pola, Vineeth N. Balasubramanian
arXiv Machine Learning
Aug 19

CORAM: Coherent Orthogonal Rotation for Model Merging

CORAM (Coherent Orthogonal Rotation for Model Merging) is a new method for combining fine‑tuned models without joint training or access to original data. It partitions each target weight matrix into row slices, represents each expert slice with its singular value decomposition in the base‑model SVD frame, and merges the task‑specific factors on their corresponding manifolds. The approach includes an amplification coefficient to counteract manifold averaging contraction, spread slicing to balance highly updated rows, and a residual pathway for non‑target layers, achieving improvements over existing orthogonal merging techniques across multiple model families and scales.

By Xinyi Sui, Ziran Liu, Nam Ling, Wei Wang, Wei Jiang
Simon Willison
Sep 4

The Pelican comparison grid for Astra is pretty interesting

Simon Willison tested GPT‑6 Astra by generating SVG pelicans riding bicycles at various reasoning levels and compared the results to GPT‑5.6 Sol, Terra, and Luna. The Astra pelicans consistently outperformed the other models, especially at low and xhigh reasoning levels, and even the Astra max version produced high‑quality images. Astra also used fewer tokens and was roughly twice as expensive as Sol, yet its low‑level output was cheaper and superior to any Sol model.

arXiv AI
Jul 20

When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis

arXiv:2607. 16062v1 Announce Type: cross Abstract: Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting this substitution is essentially never tested against the baseline it claims to replace: methods merge independently released agents precisely because a joint model is unavailable.

By S. Aaron McClendon
Hugging Face Trending Papers
Aug 18

CORAM: Coherent Orthogonal Rotation for Model Merging

CORAM introduces a new approach to merging fine‑tuned models by partitioning each target weight matrix into row slices and representing each slice with its singular value decomposition in the base‑model’s SVD frame. The method performs manifold averaging of task‑specific factors and applies an amplification coefficient to counteract contraction, with the coefficient’s scale estimated from update norms and its restoration strength chosen from expert update dispersion. Across multiple model families and scales, CORAM outperforms the prior OrthoMerge technique by up to 1.35 points and matches or exceeds the strongest weight‑space baselines.

arXiv Computation and Language
6d ago

Orthogonal Yet Coupled: Decoupling Geometric Components for Model Merging

The paper introduces DiGA, a Disentangled Geometry-Aware framework for merging pretrained models. DiGA orthogonally decomposes each task vector into components tied to distinct geometric attributes, aggregates these components independently, and then recombines them, thereby preserving each component’s geometric identity. Experiments across various models, tasks, and merging methods show that DiGA improves merged-model performance and reduces capability degradation.

By Zijing Wang, Yongkang Liu, Mingyang Wang, Ercong Nie, Mengjie Zhao, Yunpu Ma, Kang Liu, Zihan Wang, Shi Feng, Daling Wang, Hinrich Sch\"utze