The paper investigates how autoencoder parameters reflect the statistical properties of their training data. By analyzing the spectral characteristics of the parameter matrices, it shows that singular values correspond to eigenvalues of the data covariance matrix, linking data and parameter spaces. Experiments on CIFAR‑10 and FashionMNIST demonstrate that these spectral vectors can accurately differentiate models trained on different data subsets without complex generation methods or access to the original samples.
arXiv:2608.29867v1 Announce Type: new
Abstract: Autoencoders are widely used for nonlinear dimensionality reduction and manifold learning. While most common implementations rely on both nonlinear enc...
By Louen Pottier, Louis Lesueur, Anders Thorin
arXiv:2602. 10680v2 Announce Type: replace-cross Abstract: Many real-world datasets contain hidden structure that cannot be detected by simple linear correlations between input features.
By Vicente Conde Mendes, Lorenzo Bardone, C\'edric Koller, Jorge Medina Moreira, Vittorio Erba, Emanuele Troiani, Lenka Zdeborov\'a
arXiv:2606. 25900v1 Announce Type: new Abstract: Variational Autoencoders (VAEs) belong to a family of autoencoders with probabilistic properties, making them well suited for generating data by producing a smooth and continuous latent space.
By Gananath R
Variational Autoencoders (VAEs) belong to a family of autoencoders with probabilistic properties, making them well suited for generating data by producing a smooth and continuous latent space. Despite being introduced over a decade ago, the method continues to be widely adopted in both research and industry for diverse applications.
arXiv:2607. 01275v1 Announce Type: cross Abstract: Variational Autoencoders (VAEs) commonly assume a standard isotropic Gaussian prior over the latent space, an assumption that often fails to capture the true distribution of latent representations for complex datasets.
By Qijun Chen, Shaofan Li
arXiv:2501. 09876v3 Announce Type: replace-cross Abstract: Generative modeling aims to generate new data samples that resemble a given dataset.
By Wonjun Lee, Riley C. W. O'Neill, Dongmian Zou, Jeff Calder, Gilad Lerman
arXiv:2609.28409v1 Announce Type: cross
Abstract: Vector Symbolic Algebras project data structures into a hyperdimensional vector space through the application of their vector algebras to randomly ge...
By Mohamed Malek Abid, P. Michael Furlong
Introduction Heavy computation is a well-known problem in various ML algorithms today, especially when generative AI is applied to text, images, and other unstructured data. One of the principal approaches to mitigate this problem is to compress input data into a lower-dimensional representation while preserving the main context.
By Vyacheslav Efimov
The paper presents a theoretical analysis of symmetric autoencoders, a class of deep learning architectures frequently used in machine learning tasks. It distinguishes between different symmetric designs and shows that the reconstruction error of orthonormal symmetric autoencoders can be interpreted via the Eckart‑Young‑Schmidt theorem. Building on this insight, the authors propose an EYS‑based initialization strategy using repeated SVD, and validate its effectiveness through numerical experiments comparing it to conventional deep autoencoders.
By Simone Brivio, Nicola Rares Franco
The paper introduces a spectrally-aligned latent-flow model for time‑series generation that trains the latent space to preserve dynamical properties relevant to synthetic data quality. By incorporating fine‑tuning losses based on Fourier, wavelet, and signature transforms, the method mitigates spectral mismatches caused by latent compression and ensures alignment with true signals in terms of smoothness and targeted spectral content. Experiments on real‑world long‑range univariate and multivariate benchmarks show that the aligned model outperforms a base latent‑flow model and state‑of‑the‑art approaches in signal realism, computational efficiency, and local structure alignment.
By Camilo Carvajal Reyes, Felipe Tobar
arXiv:2604. 00669v2 Announce Type: replace Abstract: This study examines the challenges of modeling complex and noisy data related to socioeconomic factors over time, with a focus on data from various districts in Odisha, India.
By Sandeep Kumar Samota, Reema Gupta, Snehashish Chakraverty