arXiv Machine Learning

Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent

arXiv:2607. 22474v1 Announce Type: new Abstract: In overparameterized linear regression, many weak spectral directions act like a ridge penalty on the signal-bearing spectrum; negative ridge is the natural correction, pushing filters above one.

arXiv Machine Learning
Jul 16

Gauge-Invariant, Parameter-Insensitive Regularization for Potential Recovery from Flow on Directed Graphs

arXiv:2607. 13609v1 Announce Type: new Abstract: Recovering a latent potential from observed flow on a directed graph (a discrete Poisson problem with Dirichlet boundaries) is ill-posed, and the standard fix backfires: ridge regularization shrinks toward a gauge-meaningless origin, collapsing and reversing the recovered ordering ($+0.

By Mohammad Forouhesh
arXiv Machine Learning
Sep 1

Singular Curvature in ReLU Training:Differentiation and the Gradient-Flow Limit Need Not Commute

The paper investigates the relationship between discrete gradient descent (GD) and its continuous-time gradient-flow counterpart in the context of ReLU neural networks. It shows that while GD states converge over a finite horizon, the exact discrete derivatives obtained via automatic differentiation do not necessarily match the derivative of the limiting flow, due to singular curvature at activation events. The authors provide a Stieltjes representation that separates continuous regional Hessians from atomic interface curvature, revealing rank-one discrepancies at activation jumps and demonstrating that even globally strongly convex residual-ReLU losses can exhibit large sensitivity ratios on certain initialization sets.

By Xiaoyang Li, Runni Zhou
arXiv Machine Learning
Sep 4

Restricted Eigenvalues Beyond Gaussian Width: Threshold Occupancy under Heavy Tails

The paper investigates restricted eigenvalue (RE) bounds for norm‑regularized estimators under heavy‑tailed designs. It shows that the previously conjectured sample‑size law based on Gaussian width fails for heavy‑tailed measurements, due to a phenomenon called simultaneous threshold occupancy. The authors provide explicit counterexamples, derive worst‑case sample‑complexity bounds, and compare the behavior of heavy‑tailed versus Gaussian designs on constant‑width polyhedral descent cones.

By Shi Fu, Huibo Xu, Qixin Zhang, Dacheng Tao
arXiv AI
6d ago

Spectral-Sphere-Constrained Hyper-Connections

The paper introduces Spectral‑Sphere‑Constrained Hyper‑Connections (s²HC), a new method for controlling the residual matrices used in Hyper‑Connections (HC). Unlike previous doubly stochastic constraints that caused identity degeneration, expressivity bottlenecks, and parameterization inefficiencies, s²HC confines these matrices to a spectral norm sphere, restoring flexibility over subdominant spectra and eliminating unstable Sinkhorn‑Knopp iterations. This approach preserves training stability while allowing expressive, non‑degenerate residual matrices.

By Zhaoyi Liu, Haichuan Zhang, Ang Li
arXiv Machine Learning
Jun 10

Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization

arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.

By Jingfeng Wu, Peter L. Bartlett, Sham M. Kakade, Jason D. Lee, Bin Yu