arXiv Computer Vision
Sep 22

Bridging Reconstruction and Generation: A Latent Distribution Perspective on Evaluation and Improvement

The paper investigates why reconstruction quality in latent generative models does not always predict generative performance, attributing the issue to a mismatch between encoder-induced and generation-time latent distributions. It introduces Generation‑Aware Reconstruction (GAR), a method that perturbs encoder latents with noise and denoises them through the generative model before decoding, creating a continuous trajectory that reveals how the decoder behaves across latent spaces. The resulting GAR‑FID metric correlates strongly with generation FID, and using intermediate GAR latents for decoder adaptation consistently improves generative quality across different model scales.

By Xianghong Fang, Wenjie Shu, Tongda Xu, Wenlong Mou, Dehan Kong, Tim G. J. Rudner