🤗 PEFT welcomes new merging methods
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
The article introduces Guided Merge Sort, an optimized sorting technique that combines elements of ordinary merge sort and multi‑way merge sort. It highlights how the use of the "goto" operator becomes essential in this approach. The post explains the algorithm’s design and its potential advantages over traditional methods.
Merge++ is a post‑hoc refinement technique for model merging that synthesizes task‑representative images by inverting expert checkpoints and then distills expert knowledge into a merged model. It operates without any additional data beyond the checkpoints and can be applied universally across existing weight‑space merging algorithms. Experiments show consistent improvements, with average gains of +2 to +8 points and up to +25.9 on specific configurations.
arXiv:2607. 20729v1 Announce Type: cross Abstract: A record system declares when two records refer to the same entity, occurrence, scope, or rule.
CORAM (Coherent Orthogonal Rotation for Model Merging) is a new method for combining fine‑tuned models without joint training or access to original data. It partitions each target weight matrix into row slices, represents each expert slice with its singular value decomposition in the base‑model SVD frame, and merges the task‑specific factors on their corresponding manifolds. The approach includes an amplification coefficient to counteract manifold averaging contraction, spread slicing to balance highly updated rows, and a residual pathway for non‑target layers, achieving improvements over existing orthogonal merging techniques across multiple model families and scales.
Simon Willison tested GPT‑6 Astra by generating SVG pelicans riding bicycles at various reasoning levels and compared the results to GPT‑5.6 Sol, Terra, and Luna. The Astra pelicans consistently outperformed the other models, especially at low and xhigh reasoning levels, and even the Astra max version produced high‑quality images. Astra also used fewer tokens and was roughly twice as expensive as Sol, yet its low‑level output was cheaper and superior to any Sol model.