arXiv Machine Learning

How Neural Losses Shape VAE Latents

arXiv:2606. 00635v1 Announce Type: new Abstract: Modern VAEs are rarely trained with the pointwise likelihood implied by the standard $\beta$-VAE objective.

arXiv Machine Learning
Sep 24

On the Diffusibility of High-Dimensional Latents

The paper investigates how fine‑tuning pretrained visual encoders for faithful image reconstruction affects diffusion models that operate in the resulting latent space. It finds that such fine‑tuning reduces the effective dimensionality of the latent representation, causing standard velocity‑prediction flow‑matching to fit noise outside the low‑dimensional signal manifold and making optimization inefficient. Consequently, the authors propose using a clean‑data ($oldsymbol{x}_{0}$) parameterization, which focuses learning on the signal manifold and consistently improves text‑to‑image generation across multiple strong‑reconstruction encoders.

By Chao Feng, Zhiyang Xu, Bowei Chen, Yuanjun Xiong, Xiyao Wang, Jui-Hsien Wang, Richard Zhang, Zhe Lin, Andrew Owens, Yijun Li