$\beta$-VAEs as Effective Theories: Tolerance-Dependent Dimension
arXiv:2608. 10599v1 Announce Type: new Abstract: In a $\beta$-VAE, increasing the regularization strength acts as a spectral cutoff by collapsing low-utility latent coordinates.
arXiv:2605. 22691v2 Announce Type: replace Abstract: We show that, in linear Gaussian VAEs, posterior collapse is a form of latent feature selection.
arXiv:2608. 10599v1 Announce Type: new Abstract: In a $\beta$-VAE, increasing the regularization strength acts as a spectral cutoff by collapsing low-utility latent coordinates.
arXiv:2607. 05531v1 Announce Type: new Abstract: Variational Autoencoders (VAEs) frequently suffer from posterior collapse, a failure mode in which the approximate posterior converges to the prior, rendering the latent code uninformative.
arXiv:2509. 24882v2 Announce Type: replace Abstract: Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models.
arXiv:2608. 09417v1 Announce Type: new Abstract: Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate under conventional initialization schemes.
arXiv:2607. 26414v1 Announce Type: cross Abstract: Latent low-dimensional structure in datasets of natural and engineered systems enables their sparse sensing, or full-state reconstruction from historical data and very few carefully chosen localized measurements.
arXiv:2606. 04405v1 Announce Type: cross Abstract: Modern Transformer architectures frequently employ normalization mechanisms such as RMSNorm and Query-Key Normalization, making parts of the model approximately scale-invariant with respect to weight magnitudes.
arXiv:2511. 11927v2 Announce Type: replace-cross Abstract: Principal Component Analysis (PCA) is a standard tool for extracting a low-rank signal from noisy observations.
arXiv:2608. 09417v2 Announce Type: replace Abstract: Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate under conventional initialization schemes.
arXiv:2606. 06576v1 Announce Type: new Abstract: In the sciences, regression tasks often require predicting high-dimensional outputs from few training examples.
arXiv:2608. 11321v1 Announce Type: cross Abstract: We study spectral clustering in the presence of a confounding latent geometry.
arXiv:2606. 14533v1 Announce Type: new Abstract: Principal Component Analysis (PCA) preserves variance, not the information needed to detect rare catastrophic events.
arXiv:2606. 29723v1 Announce Type: new Abstract: Continuous physical fields represent a large fraction of data under scientific investigation.