The paper introduces a Projected Riemannian Gradient Descent (RGD) algorithm for computing the Bures‑Wasserstein barycenter of positive definite matrices, achieving dimension‑independent linear convergence at unit step size. It resolves a previous dichotomy by showing that clipping eigenvalues to a fixed interval yields a closed‑form, non‑expansive projection in the BW metric, allowing the algorithm to match the empirical speed of unit‑step RGD while maintaining theoretical guarantees. The method also extends to the invariant matrix projection problem, providing a unified dimension‑independent analysis.
arXiv:2609. 03762v1 Announce Type: new Abstract: The computation of the Bures-Wasserstein (BW) barycenter of an ensemble of positive definite matrices arises throughout machine learning, optimal transport, and quantum information.
By A. Afham
arXiv:2608. 15982v1 Announce Type: new Abstract: We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces.
By Mahdi Mohammadigohari, Thomas Borsani, Giuseppe Di Fatta
We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces. In vector-valued Sobolev RKHSs, we derive Rademacher complexity bounds for invertible and width-expanding injective architectures.
arXiv:2608. 13882v1 Announce Type: new Abstract: Claims about the benefit of depth depend on the complexity assigned to a representation.
By Mahdi Mohammadigohari
arXiv:2606. 01443v1 Announce Type: cross Abstract: A central difficulty in training Joint-Embedding Predictive Architectures (JEPAs) is preventing representation collapse.
By Triet M. Le
The paper introduces a new representation‑adaptive kernel class that, on a fixed sample, yields a union of reproducing‑kernel Hilbert‑space ellipsoids instead of a single ellipsoid. It defines a minimum‑trace common covariance dominating the empirical union generated by Brownian kernel ladders, and derives exact formulations, statistical and computational consequences, and a universal Gaussian‑complexity bound. The work further develops geometric reductions, deterministic depth laws, and exact empirical Kolmogorov‑width formulas, providing both lower and upper certificates for covariance certification and illustrating the distinction between successful covariance certification and predictive selection.
By Mahdi Mohammadigohari
arXiv:2510. 22778v3 Announce Type: replace-cross Abstract: We develop a free-probabilistic framework for denoising diffusion, in which the data is a self-adjoint operator and its law a spectral distribution.
By Swagatam Das
The paper studies operator learning on function spaces using encoder–decoder architectures. It shows that as input and output resolutions grow, the induced kernels converge to a limiting kernel, enabling regularity assumptions independent of resolution. The authors derive upper and lower bounds for regularized stochastic gradient descent, extend the analysis to neural networks via the limiting neural tangent kernel, and provide error bounds and complexity guarantees for various kernel and encoding constructions.
By Lei Shi, Jia-Qi Yang, Ding-Xuan Zhou
arXiv:2605. 08170v2 Announce Type: replace Abstract: Neural operators have emerged as a powerful tool for learning mappings between infinite-dimensional function spaces.
By Nicole Hao
arXiv:2606. 30523v1 Announce Type: new Abstract: Covariance matrices serve as compact descriptors of feature distributions in many machine-learning pipelines, including domain adaptation and Gaussian embeddings.
By Woojoo Na, Jennifer Dy
arXiv:2606. 24157v1 Announce Type: new Abstract: The space $\mathcal{P}_2(\mathbb{R}^d$) of probability measures with finite second moment carries a natural geometry: the quadratic Wasserstein distance W_2 makes it a complete metric space and, following Otto, a (formal) Riemannian manifold whose geodesics are the optimal-transport interpolations.
By Yian Yao, Weiwei Zhang