arXiv AI By Chenrui Wang, Yixuan Qiu

Accelerating Birkhoff Projection for Manifold-Constrained Hyper-Connections

Read the original on arXiv AI →

arXiv:2606. 07574v1 Announce Type: cross Abstract: Manifold-constrained hyper-connections (mHCs) have recently been proposed as a principled extension of hyper-connections, where the residual mixing matrices are constrained to be doubly stochastic via projection onto the Birkhoff polytope.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Spectral-Sphere-Constrained Hyper-Connections

The paper introduces Spectral‑Sphere‑Constrained Hyper‑Connections (s²HC), a new method for controlling the residual matrices used in Hyper‑Connections (HC). Unlike previous doubly stochastic constraints that caused identity degeneration, expressivity bottlenecks, and parameterization inefficiencies, s²HC confines these matrices to a spectral norm sphere, restoring flexibility over subdominant spectra and eliminating unstable Sinkhorn‑Knopp iterations. This approach preserves training stability while allowing expressive, non‑degenerate residual matrices.

By Zhaoyi Liu, Haichuan Zhang, Ang Li
arXiv Machine Learning
Jun 15

Scalable Deep Unfolding of Conic Optimizers

arXiv:2606. 13825v1 Announce Type: cross Abstract: Deep unfolding (DU) accelerates iterative optimizers by introducing learnable components and training them through unrolled iterations, but extending DU to the large-scale semidefinite programs (SDPs) common in robotics has remained limited.

By Alex Oshin, Rahul Vodeb Ghosh, Evangelos A. Theodorou
arXiv Machine Learning
1d ago

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

The paper introduces TACO, a new optimizer for fine‑tuning large language models that drastically reduces optimizer state memory while preserving first‑order gradients. TACO selects the sign of the largest magnitude entry in each column of weight matrices, achieving a 174× reduction in persistent optimizer memory compared to AdamW8bit and a 2.9× decrease in peak training memory on OPT‑13B. This allows full‑parameter fine‑tuning of 30–32B‑parameter models on a single 80 GB GPU across multiple model families and tasks, with comparable accuracy and runtime to existing methods.

By Jichao Jiang (University of Central Florida), Cristian McGee (University of Central Florida), El Houcine Bergou (Mohammed VI Polytechnic University), Hanqin Cai (University of Central Florida), Aritra Dutta (University of Central Florida)