arXiv Machine Learning

Adaptive Runge-Kutta Step Control Buys Training Loss, Not Generalization: An Honest Compute-Matched Study of RK-Adam Optimizers

arXiv:2607. 14516v1 Announce Type: new Abstract: Interpreting optimizers as gradient-flow discretizations has motivated applying higher-order Runge-Kutta (RK) integrators to neural networks.

arXiv Machine Learning
Sep 24

The Drift Contract: Spectral Updates for Depth-Robust Local Learning

The paper introduces the Drift Contract, a spectral update geometry for local learning that improves depth robustness and hyperparameter stability. By applying momentum orthogonalization with spectral step scaling to per‑layer updates, the authors achieve consistent performance across a wide range of widths and depths on CIFAR‑10 MLPs, outperforming local Adam and providing a per‑layer, input‑conditioned drift bound. The study also shows that the spectral geometry itself, rather than step‑size rules, drives the observed depth robustness, while a negative result indicates that the stability benefit is limited to non‑normalized layers.

By Fabien Polly
arXiv Machine Learning
Sep 2

Gradient-Update Mismatch: Rethinking Conflict-Free Training of Physics-Informed Neural Networks

The paper identifies a problem called Gradient-Update Mismatch (GUM), where optimizers can alter conflict-free gradient directions produced by gradient surgery, leading to conflicts between physics residual and boundary condition losses in Physics-Informed Neural Networks (PINNs). To address this, the authors propose Gradient-Update Alignment (GUA), which projects the optimizer’s update onto the conflict-free cone and adjusts internal optimizer state accordingly. Experiments show GUM is common across many optimizers, and GUA consistently eliminates conflicts and significantly reduces error in PINN training.

By Jing Xiao, Xinhai Chen, Qinglin Wang, Menghan Jia, Zhiquan Lai, Dongsheng Li, Jie Liu, Tiejun Li
arXiv Machine Learning
Aug 4

An Embedded RISC-V Evaluation of Kolmogorov--Arnold Networks in Hard-Constrained Recurrent Physics-Informed Models

arXiv:2608. 00737v1 Announce Type: new Abstract: Hard-constrained recurrent physics-informed networks (HRPINNs) embed known dynamics inside a recurrent numerical integrator and restrict a neural branch to learning only the residual dynamics that the first-principles model does not capture.

By Enzo Nicolas Spotorno, Josafat Leal Filho
arXiv Machine Learning
Sep 4

Activation-Keyed Momentum: An Anisotropic Momentum Update via the Delta Rule

The paper introduces Activation-Keyed Momentum (AK‑Momentum), a momentum update that uses the input activation of a linear layer as a key to apply a delta‑rule update, allowing each direction to decay at a rate proportional to its frequency of appearance. AK‑Momentum is proven to be a valid momentum, incorporates input‑side curvature correction without matrix inversion, and clears stale directions faster than traditional exponential moving average (EMA) under both fixed and drifting optima. It can replace the momentum buffer of any optimizer, scales with width under μP, adds only 22–25% extra compute, and demonstrates significant step‑count reductions in FineWeb‑Edu pretraining and other benchmarks. whyItMatters":"AK‑Momentum offers a principled, efficient way to adapt momentum decay to anisotropic training dynamics, improving convergence speed and stability across a range of models and datasets."

By Euijin Hong, Guannan Qu
arXiv AI
Jul 8

K-ABENA: K-Adaptive Backpropagation with Error-based N-exclusion Algorithm : (Compensated Loss-Based Sample Exclusion with Unbiased Gradient Estimation)

arXiv:2607. 05903v1 Announce Type: cross Abstract: We present K-ABENA (K-Adaptive Backpropagation with Error-based N-exclusion Algorithm), a selective gradient computation framework that reduces per-iteration training cost by excluding a fraction of low-loss ("minor") observations from the backward pass.

By Jean-Francois Bonbhel
arXiv Machine Learning
Jul 27

Hyperball May Not Be a Free Lunch

arXiv:2607. 22444v1 Announce Type: new Abstract: For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing updates.

By Yihao Xiao, Jialong Sun, Zitian Gao, Zeming Wei, Chutian Wang, Ran Tao, Jiaye Teng, Bryan Dai
arXiv AI
Sep 25

Wiring Beats Blending: Structure-Aware Compensation for Transformer Downscaling

The paper investigates converting a large pretrained transformer (1.4 B parameters) into a smaller sibling (410 M) by studying representation alignment and parameter projection. It finds that dense weight projection destroys structure, and that a low‑budget, structure‑aware compensation—separating least‑squares function alignment from variance‑preserving rescaling—yields significant gains on token‑efficient training, outperforming subcloning and standard distillation pipelines at matched budgets.

By Ravi Satya Durga Prasad Yenugula