The paper introduces Spectral‑Sphere‑Constrained Hyper‑Connections (s²HC), a new method for controlling the residual matrices used in Hyper‑Connections (HC). Unlike previous doubly stochastic constraints that caused identity degeneration, expressivity bottlenecks, and parameterization inefficiencies, s²HC confines these matrices to a spectral norm sphere, restoring flexibility over subdominant spectra and eliminating unstable Sinkhorn‑Knopp iterations. This approach preserves training stability while allowing expressive, non‑degenerate residual matrices.
By Zhaoyi Liu, Haichuan Zhang, Ang Li
arXiv:2606. 02596v1 Announce Type: new Abstract: The curvature exponent $\alpha$ in $h_k \propto \sigma_k^\alpha$ -- governing how Hessian eigenvalues scale with gradient singular values -- varies systematically across layer types ($\alpha \approx 2$ for convolutions, $\approx 1$ for transformer attention, $< 1$ for MLP up-projections).
By Anherutowa Calvo
The paper introduces Spectrally Optimised Neural Discretisations (SpeND), a mesh‑free framework that learns discretisation weights from local stencil geometry on unstructured point clouds. By embedding discrete moment conditions into the network architecture, SpeND guarantees polynomial consistency and allows the weights to be optimised for spectral accuracy over a chosen wavenumber band, using an unsupervised Fourier‑mode loss. The resulting operators are PDE‑agnostic, perform well on Poisson, Burgers, and Navier–Stokes equations, and can reduce wall‑clock time by 3–20× compared to existing mesh‑free methods at the same accuracy.
By Lucas Gerken Starepravo, Henry Broadley, Steven Lind, Jack R. C. King
SPARCL introduces a spectral partitioned analytic continual learning method that addresses forgetting in analytic class‑incremental learning. By decomposing the running autocorrelation into a high‑energy core and a residual complement, SPARCL freezes core components for old classes and updates only the residual block, ensuring closed‑form updates with an invariance guarantee. Experiments on CIFAR‑100, CUB‑200, ImageNet‑R, and ImageNet‑A with a frozen ViT‑B/16 protocol show that SPARCL narrows the performance gap between classical analytic learners and strong representation matchers while complementing sparse feature‑decorrelation approaches.
By James Hartley, Zeropy Surio, Daniel Whitmore, Hannah Clarke, Thomas Reed
arXiv:2607. 07032v2 Announce Type: replace Abstract: Spectral positional encodings (PEs) for \emph{directed} graphs face two obstacles: magnetic Laplacians require an $O(n^3)$ Hermitian eigendecomposition per potential, and their complex eigenvectors are defined only up to unitary gauge, which prior work handles with basis-invariant architectures.
By Jiaqing Xie, Yuxin Wang
arXiv:2606. 28444v1 Announce Type: cross Abstract: Classical universal approximation theorems establish the expressive power of sigmoidal multilayer perceptrons, but they do not prescribe how initial weights should encode the geometry of a data distribution.
By Yi-Shan Chu
arXiv:2602. 00722v2 Announce Type: replace Abstract: Parameter-efficient continual learning aims to adapt pre-trained models to sequential tasks without forgetting previously acquired knowledge.
By Hao Gu, Mao-Lin Luo, Zi-Hao Zhou, Han-Chen Zhang, Min-Ling Zhang, Tong Wei
arXiv:2607. 22931v1 Announce Type: new Abstract: Analytic Continual Learning (ACL) offers a computationally efficient alternative to gradient-based approaches.
By Quyen Tran, Hai Nguyen, Quan Dao, Zhuowei Li, Nam Le, Trung Le, Dimitris Metaxas
arXiv:2607. 07032v1 Announce Type: new Abstract: Spectral positional encodings (PEs) for \emph{directed} graphs face two obstacles: magnetic Laplacians require an $O(n^3)$ Hermitian eigendecomposition per potential, and their complex eigenvectors are defined only up to unitary gauge, which prior work handles with basis-invariant architectures.
By Jiaqing Xie, Yuxin Wang
arXiv:2608.19021v2 Announce Type: replace
Abstract: Global Covariance Pooling (GCP) improves deep networks by capturing second-order feature statistics, and is especially effective for fine-grained r...
By Md Rifat Ur Rahman, Md Raihan Khan, Md Sakib Hossain Shovon, Pietro Li\`o, Mohammad Ali Moni
arXiv:2609.21039v1 Announce Type: new
Abstract: A pervasive structural pattern in modern deep learning is the linear factorization block: a submodule of the form $W = BA$ in which two parameter matri...
By Emanuele Zangrando, Marco Sutti, Francesco Tudisco
The paper introduces a learning-oriented framework for spectral certification, sensitivity analysis, and adaptive control of directed graph learning models using cone extended Rayleigh quotients. It provides computable cone bounds and differentiable soft-min/max surrogates that enable rigorous one-sided spectral bounds without requiring symmetry or cone preservation. Experiments on directed networks, including the Cora citation graph, demonstrate that adaptive sensitivity recomputation can significantly reduce spectral levels while preserving test accuracy.
By Yavdat Sh. Il'yasov, Nur F. Valeev