The paper introduces the spherical Cauchy distribution as a new hyperspherical posterior for variational autoencoders, avoiding the complications of the von Mises–Fisher and Power Spherical alternatives. By using stereographic projection and a Möbius transformation, the authors obtain exact posterior samples and a closed‑form KL divergence that terminates in a finite polynomial for even dimensions and admits certified truncation for odd dimensions. Empirical results show that the spherical Cauchy yields faster inference and lower reconstruction loss on MNIST and improved negative log‑likelihood on smallNORB compared to existing methods.
By Lukas Sablica, Kurt Hornik
arXiv:2609.25659v1 Announce Type: new
Abstract: Many scientific datasets, such as molecular conformational ensembles or single-cell tissue measurements, are naturally modeled as meta-distributions: d...
By Doron Haviv, Edward De Brouwer, Rishabh Anand, Rex Ying, A\"icha Bentaieb, Gabriele Scalia, Hector Corrada Bravo
arXiv:2605. 05629v3 Announce Type: replace-cross Abstract: We study the problem of learning generative models for discrete sequences in a continuous embedding space.
By Jannis Chemseddine, Gregor Kornhardt, Gabriele Steidl
arXiv:2506. 21278v3 Announce Type: replace-cross Abstract: We propose spherical Cauchy (spCauchy) latent variables for variational autoencoders on hyperspherical latent spaces.
By Lukas Sablica, Kurt Hornik
arXiv:2607. 03329v1 Announce Type: new Abstract: Conventional uniform convergence bounds and empirical risk minimization break down in massive over-parameterized models, such as large language transformers and biological sequence networks.
By Bing Cheng, Yi-Shuai Niu, Howell Tong, Shing-Tung Yau
arXiv:2404.17763v3 Announce Type: replace-cross
Abstract: Probabilistic graphical models that encode an underlying Markov random field are fundamental building blocks of generative modeling to learn...
By Yujie Chen, Anindya Bhadra, Antik Chakraborty
The paper extends the analysis of Joint-Embedding Predictive Architectures (JEPAs) beyond Euclidean latent spaces to Riemannian manifolds. It shows that when latent variables lie on a sphere and the target distribution matches this spherical geometry, every optimal representation recovers the latent state up to an orthogonal transformation, demonstrating that Gaussian uniqueness is not universal. Experiments confirm that geometrically compatible targets improve linear recovery, especially in high-dimensional toroidal settings.
By L\'eo Nicollier (CB, ATT), Enric Meinhardt-Llopis (CB), Marc Pic (ATT), Pablo Mus\'e (CB, IFUMI), Gabriele Facciolo (CB)
arXiv:2505. 04338v3 Announce Type: replace Abstract: We propose Riemannian Denoising Diffusion Probabilistic Models (RDDPMs) for learning distributions on submanifolds of Euclidean space that are level sets of functions, including most of the manifolds relevant to applications.
By Zichen Liu, Wei Zhang, Christof Sch\"utte, Tiejun Li
arXiv:2606. 07058v1 Announce Type: new Abstract: Variational autoencoders (VAEs) learn low-dimensional latent representations of high-dimensional data.
By Jilles S. van Hulst, Jakub M. Tomczak, W. P. M. H. Heemels, Duarte J. Antunes
arXiv:2502. 14424v3 Announce Type: replace-cross Abstract: Most self-supervised learning objectives defend against collapse but leave the target representation law unspecified.
By Yuling Jiao, Wensen Ma, Defeng Sun, Hansheng Wang, Yang Wang
arXiv:2609.10305v1 Announce Type: new
Abstract: Language models under one million parameters matter for edge deployment, domain adaptation, and reproducible research, yet a two-layer LSTM or Transfor...
By Fang Li
arXiv:2606. 15760v1 Announce Type: new Abstract: A significant gap exists between theory and practice in deep learning.
By Marios Koulakis, Constantin Seibold