arXiv Machine Learning

DREG: A Layer-Wise Jacobian Regularization as a General-Purpose Penalty

arXiv:2606. 23942v1 Announce Type: new Abstract: We present a large-scale empirical study isolating the contributions of the Derivative Regularization penalty (DREG).

arXiv Machine Learning
Sep 23

Double Descent and Malign Overfitting in Diffusion Models

The paper investigates why diffusion models, unlike typical deep learning models, exhibit catastrophic overfitting when overparameterized. Through experiments on U‑Nets trained on CelebA and a random‑features theoretical analysis, it shows that the interpolation peak occurs at a model size proportional to the product of training samples and noise realizations, but the test loss starts to rise already at the number of samples, leading to memorization of the empirical score. Regularization techniques such as ridge penalties or early stopping can still make large models outperform smaller, unregularized ones.

By Rapha\"el Urfin, Tony Bonnaire, Giulio Biroli, Marc M\'ezard
Hugging Face Trending Papers
Jul 15

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip connections and normalization, long understood as controlling magnitude, as mechanisms for preserving gradient rank across depth, since the very matrix multiplications and nonlinear activations that make the network expressive also reduce the rank.

arXiv Computer Vision
Sep 15

Sparsity-Adaptive Sharpness-Aware Minimization

The paper introduces Sparsity-Adaptive Sharpness-Aware Minimization (SA‑SAM), a method that adjusts the perturbation radius in sharpness-aware training to remain consistent as model sparsity increases. It also evaluates a Magnitude‑Weighted Hessian (MWH) importance metric derived from second‑order analysis. Experiments on CIFAR‑10‑C, CIFAR‑100‑C, and ImageNet‑100‑C show that SA‑SAM improves corruption robustness at 80–90% sparsity while maintaining clean accuracy, and the study reports inference throughput at deployment‑relevant sparsity levels.

By Shiryu Ueno, Yoshikazu Hayashi, Kunihito Kato