arXiv:2606. 16214v1 Announce Type: cross Abstract: Modern deep learning models remain notoriously prone to overconfidence, limiting their reliability in high-stakes applications.
By Tobias Jan Wieczorek, Leon de Andrade, Thomas M\"ollenhoff, Marcus Rohrbach
The paper investigates why adaptive optimizers like Adam outperform SGD when fine‑tuning Transformers. It introduces gradient heterogeneity—the variation in gradient norms across parameter blocks—and shows, both theoretically and experimentally, that this heterogeneity, together with Hessian heterogeneity, hampers SGD convergence while sign‑based methods such as SignSGD are less affected. The study links the source of gradient heterogeneity to layer‑normalization placement, finding that Post‑LN architectures exhibit the strongest effect, and uses SignSGD as a tractable proxy to analyze Adam‑like behavior and learning‑rate scaling.
By Akiyoshi Tomihari, Issei Sato
arXiv:2607. 10593v1 Announce Type: new Abstract: Normalization is a critical component for stabilizing Transformer training, yet the choice between static strategies such as Layer Normalization (LN) and adaptive alternatives remains largely task-dependent.
By Piyush Kaushik Bhattacharyya, Divyanshu Rai, Swastik Singh, Kumar Aakash, Ayush Ranjan, Krutika Verma
arXiv:2601. 22580v2 Announce Type: replace-cross Abstract: The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures.
By Chao Wang, Bei Li, Jiaqi Zhang, Xinyu Liu, Yuchun Fan, Linkun Lyu, Xin Chen, Jingang Wang, Tong Xiao, Peng Pei, Xunliang Cai
arXiv:2509. 25136v3 Announce Type: replace Abstract: Activation-aware low-rank factorization techniques yield strong compression results but are generally confined to linear layers, while existing whitening-based theory typically makes an implicit full-rank assumption on activations.
By David Gonz\'alez-Mart\'inez
The paper introduces Anon, an optimizer that extends adaptivity beyond the traditional bounds of SGD and Adam by allowing extrapolation across the entire real-number spectrum. It addresses the limitations of prior tunable optimizers that only interpolate between 0 and 1 adaptivity, showing that optimal adaptivity can require negative values for CNNs or values greater than one for Transformers. Anon incorporates Incremental Delay Update (IDU) to maintain provable stability and demonstrates competitive performance on image classification, diffusion, and large language modeling tasks.
By Yiheng Zhang, Kaiyan Zhao, Shaowu Wu, Yiming Wang, Jiajun Wu, Leong Hou U, Steve Drew, Xiaoguang Niu
arXiv:2505.16157v3 Announce Type: replace
Abstract: Transformer-based models have made remarkable progress in image restoration (IR) tasks. However, the quadratic complexity of self-attention in Tran...
By Yuang Ai
arXiv:2603.11323v2 Announce Type: replace
Abstract: The simplicity and effectiveness of UNet architectures make them ubiquitous in image restoration, segmentation, and diffusion models. They are ofte...
By J\'er\'emy Scanvic, Quentin Barth\'elemy, Juli\'an Tachella
arXiv:2607. 14466v1 Announce Type: new Abstract: Noise injection is a well-known technique in stochastic optimization.
By Matt L. Wiemann, Peter Melchior, Andrew K. Saydjari
The paper introduces a hardware‑aware framework that uses genetic programming to evolve layer‑specific scalar functions for Vision Transformers, replacing traditional LayerNorm with efficient, heterogeneous approximations. By applying a post‑training re‑alignment strategy, the method eliminates the need for full model retraining while achieving 90‑93% variance capture and recovering over 84% of ImageNet‑1K Top‑1 accuracy for ViT‑B and ViT‑L. The resulting architecture removes the global reduction bottleneck, reducing arithmetic complexity and off‑chip memory traffic, thereby enabling efficient deployment of ViTs on edge accelerators.
By Kieran Carrigg, Sigur de Vries, Amirhossein Sadough, Marcel van Gerven
arXiv:2609.14690v1 Announce Type: new
Abstract: Onboard satellites must restore a channel-degraded image on a few watts, using neuromorphic accelerators (e.g., BrainChip Akida, Intel Loihi-2) that su...
By Thanh-Dung Le, Vu Nguyen Ha, Ti Ti Nguyen, Symeon Chatzinotas
arXiv:2606. 02267v1 Announce Type: new Abstract: The vulnerability of deep neural networks to adversarial examples poses a significant challenge for real-world deployment.
By Nicolas Stalder, Benjamin F. Grewe, Matteo Saponati, Pau Vilimelis Aceituno