arXiv:2601. 21579v2 Announce Type: replace-cross Abstract: The success of Hyper-Connections (HC) in neural networks (NN) has also highlighted issues related to training instability and restricted scalability.
By Wuyang Zhou, Yuxuan Gu, Giorgos Iacovides, Danilo Mandic
The paper introduces Spectral‑Sphere‑Constrained Hyper‑Connections (s²HC), a new method for controlling the residual matrices used in Hyper‑Connections (HC). Unlike previous doubly stochastic constraints that caused identity degeneration, expressivity bottlenecks, and parameterization inefficiencies, s²HC confines these matrices to a spectral norm sphere, restoring flexibility over subdominant spectra and eliminating unstable Sinkhorn‑Knopp iterations. This approach preserves training stability while allowing expressive, non‑degenerate residual matrices.
By Zhaoyi Liu, Haichuan Zhang, Ang Li
arXiv:2606. 13825v1 Announce Type: cross Abstract: Deep unfolding (DU) accelerates iterative optimizers by introducing learnable components and training them through unrolled iterations, but extending DU to the large-scale semidefinite programs (SDPs) common in robotics has remained limited.
By Alex Oshin, Rahul Vodeb Ghosh, Evangelos A. Theodorou
arXiv:2609. 08136v1 Announce Type: new Abstract: This paper introduces rlaopt, a PyTorch-based package for large-scale optimization and scientific computing using randomized numerical linear algebra (RandNLA).
By Pratik Rathore, Zachary Frangella, Parth Nobel, Xuning Hu, Madeleine Udell
arXiv:2601. 02451v2 Announce Type: replace-cross Abstract: Graph Neural Networks (GNNs) suffer from over-smoothing in deep architectures and expressiveness bounded by the 1-Weisfeiler-Leman (1-WL) test.
By Subhankar Mishra
The paper introduces TACO, a new optimizer for fine‑tuning large language models that drastically reduces optimizer state memory while preserving first‑order gradients. TACO selects the sign of the largest magnitude entry in each column of weight matrices, achieving a 174× reduction in persistent optimizer memory compared to AdamW8bit and a 2.9× decrease in peak training memory on OPT‑13B. This allows full‑parameter fine‑tuning of 30–32B‑parameter models on a single 80 GB GPU across multiple model families and tasks, with comparable accuracy and runtime to existing methods.
By Jichao Jiang (University of Central Florida), Cristian McGee (University of Central Florida), El Houcine Bergou (Mohammed VI Polytechnic University), Hanqin Cai (University of Central Florida), Aritra Dutta (University of Central Florida)
arXiv:2606. 08638v1 Announce Type: cross Abstract: Recent research has developed practical, parallelizable first-order methods for large scale linear programming, but performance is highly dependent on hyperparameter selection.
By Siddharth Prasad, Dravyansh Sharma
arXiv:2607. 00095v1 Announce Type: cross Abstract: Generative models have emerged as scalable surrogates for physical simulation, yet they offer no guarantee that their outputs respect the conservation laws, boundary conditions, and nonlinear invariants that govern the underlying physics.
By Alaina Kolli, Theodoros Xenakis, Utkarsh Utkarsh, Pengfei Cai, Rafael Gomez-Bombarelli, Alan Edelman, Christopher Vincent Rackauckas
arXiv:2606. 12120v1 Announce Type: new Abstract: Low-rank optimal transport (OT) mitigates the quadratic scaling of classical solvers, yet existing approaches rely heavily on first-order mirror-descent updates that require careful hyperparameter tuning and ignore the optimization landscape's curvature.
By Pratik Jawanpuria, Bamdev Mishra
arXiv:2607. 22004v1 Announce Type: new Abstract: Energy natural gradient descent (ENGD) aligns parameter updates with the curvature of an underlying function-space energy, but existing formulations assume an unconstrained Euclidean parameter domain.
By Zhangyong Liang, Huanhuan Gao
The paper introduces HELLO, a hierarchical solver for large‑scale discrete optimal transport that reduces the problem to edge localization guided by dual potentials. HELLO uses a coarse‑to‑fine initialization across a recursive subsampling hierarchy and a refinement step that inserts the largest dual violators until a KKT residual tolerance is met, achieving linear memory usage. Experiments show that HELLO outperforms strong baselines by an order of magnitude in runtime while attaining lower transport objectives, and it scales to over a million samples in high‑dimensional settings, supporting various OT variants.
By Wenzhou Xia, Qiaoqiao Ding, Jingwei Liang, Xiaoqun Zhang
arXiv:2609.38095v1 Announce Type: new
Abstract: Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU...
By Francois Chaubard, Mykel J. Kochenderfer, Chris R\'e