arXiv AI

A Simplex Witness Certificate and Escape Force for Constant Collapse in Variational Autoencoders

arXiv:2605. 18224v4 Announce Type: replace-cross Abstract: We study exact constant collapse in variational autoencoders: the deterministic encoder mean becomes independent of the input.

arXiv AI
Aug 19

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that replaces the Gaussian regularizer used in Joint‑Embedding Predictive Architectures (JEPAs) with a contrastive inverse‑dynamics head. AC‑MTM trains a forward latent‑prediction model while an auxiliary inverse‑dynamics task forces the encoder to distinguish actions from latent transitions, preventing collapse without requiring a target network or reconstruction loss. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM matches or surpasses the performance of the Gaussian‑based SIGReg regularizer, achieving up to 20–24 point improvements on the OGBench Visual Scene benchmark.

By Jack Boylan, Chris Hokamp
arXiv Machine Learning
Aug 20

Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training

The study measured the impact of a single training example on a GPT‑2 model by running 24 counterfactual experiments. 32 models were trained from scratch on OpenWebText, and at a specific training step a single batch row was replaced with a 194‑token passage under three conditions (fluent prose, fabricated subject, random characters) or left unchanged. Results showed that the passage was learned from one exposure and decayed, with measurable differences in cross‑entropy up to 50 steps after injection but no lasting effect at the final step.

By Zachary Speck, Asa Shepard
arXiv AI
Jun 3

Exact equivariance, kept through training, buys zero-shot generalisation across the symmetry group

arXiv:2606. 03003v1 Announce Type: cross Abstract: A latent world model built from an equivariant encoder $E$ and an equivariant predictor $f$ inherits a provable symmetry of its training loss: when the world's dynamics genuinely carries a group $G$ acting on latents by an orthogonal representation $\rho(g)$, the one-step prediction relMSE is exactly invariant across the whole group, so fitting the dynamics on a restricted slice of orientations mathematically determines it on the entire orbit (j\v{u} y\=i f\v{a}n s\=an).

By Hongbo Wang (Stony Brook University)
Hugging Face Trending Papers
Aug 18

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that stabilizes Joint‑Embedding Predictive Architectures (JEPAs) without relying on Gaussian regularization. AC‑MTM adds a training‑only inverse‑dynamics head that uses Action‑NCE to force each latent transition to identify its generating action, thereby preventing encoder collapse. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM trains stably from scratch and matches or surpasses the performance of SIGReg, achieving up to a 24‑point improvement on the OGBench Visual Scene benchmark.

arXiv AI
6d ago

Latent Generative Solvers for Generalizable Long-Term Physics Simulation

The paper introduces the Latent Generative Solver (LGS), a neural PDE solver that combines a Physics VAE, a Pyramidal Flow-Forcing Transformer, and input noising to achieve generalization across twelve PDE families and stable long-term rollouts. LGS matches or surpasses deterministic baselines on one-step predictions, outperforms them on 5- and 10-step rollouts, and significantly reduces long-horizon error while cutting compute costs. It also adapts efficiently to unseen higher-resolution systems, demonstrating strong empirical performance on 2D regular-grid PDE simulations.

By Zituo Chen, Sili Deng
arXiv Computer Vision
6d ago

Conditional Predictive Sufficient Statistics for Visual Representation Learning

The paper introduces Conditional Predictive Sufficient Statistics (CPSS) as a formal way to capture useful visual representations that preserve latent factors shared with future data while discarding noise. It shows that predicting the next image patch embedding with a cosine loss approximates maximum likelihood under a von Mises-Fisher model, and that stop‑gradient alone does not enforce sufficiency. Experiments on MNIST and CIFAR‑10 with small causal Transformers demonstrate that CPSS readouts outperform intermediate blocks, while removing stop‑gradient collapses embedding rank even when the pretext loss appears perfect.

By Yuzhou Hong
arXiv AI
Sep 25

Generalized Graph Variational Autoencoders: Bounded Divergences Control Posterior Collapse

The paper introduces the Generalized Graph Variational Autoencoder (GGVA), which replaces the Kullback–Leibler divergence in the standard variational graph autoencoder with any member of the Rényi–Tsallis family of order $q$. The authors show that for $q<1$ the Tsallis divergence is bounded, whereas the KL and Rényi divergences are unbounded, and that this boundedness can significantly increase the amount of posterior information retained—up to 49× more than the VGAE on several benchmark graphs. Experiments demonstrate that the GGVA’s retained information improves node classification performance, though it does not improve link‑prediction accuracy and only delays, rather than prevents, posterior collapse.

By Kleyton da Costa, Bernardo Modenesi, Ivan F. M. Menezes, Helio Lopes
arXiv Machine Learning
Sep 25

LastOPD: Taming Collapse in Latent On-Policy Distillation

The paper introduces LastOPD, a method that mitigates collapse in latent on‑policy distillation by applying latent supervision only to the last‑layer state during a brief cross‑fade into token‑level OPD. Experiments show that LastOPD improves MATH‑500 accuracy by 5.55 and 4.02 points over token‑only OPD when distilling Qwen3‑4B and Qwen3‑8B into Qwen3‑1.7B‑Base, and achieves comparable final scores in roughly half the training steps.

By Jie Yang, Zhengyu Fang, Zelin Xu, Jiarui Sun, Xiran Fan, Junpeng Wang, Liang Wang, Qinghua Liu, Yiwei Cai, Yan Zheng