arXiv Computer Vision

Bridging Reconstruction and Generation: A Latent Distribution Perspective on Evaluation and Improvement

The paper investigates why reconstruction quality in latent generative models does not always predict generative performance, attributing the issue to a mismatch between encoder-induced and generation-time latent distributions. It introduces Generation‑Aware Reconstruction (GAR), a method that perturbs encoder latents with noise and denoises them through the generative model before decoding, creating a continuous trajectory that reveals how the decoder behaves across latent spaces. The resulting GAR‑FID metric correlates strongly with generation FID, and using intermediate GAR latents for decoder adaptation consistently improves generative quality across different model scales.

arXiv Computer Vision
Sep 3

Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies

The paper introduces Latent Drift, a generative forecasting framework that predicts slow-evolving neurodegenerative disease progression by learning changes in a compressed semantic representation rather than full-resolution anatomy. It addresses two failure modes—identity collapse and continuous interpolation trap—by removing pixel-level identity from the prediction target and applying Finite Scalar Quantization to suppress high-frequency nuisance fluctuations. Experiments on longitudinal 3D brain MRI demonstrate that Latent Drift outperforms diffusion and autoregressive transformer baselines in both generative fidelity and clinically relevant metrics.

By Yuxiang Feng, Juncheng Wang, Chao Xu, Wenlong Hou, Huihan Wang, Yijie Qian, Yang Liu, Baigui Sun, Yong Liu, Shujun Wang
arXiv AI
Aug 12

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".

By Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit
arXiv Machine Learning
4d ago

HALO: Enhancing Time Series Generation via Hyperspherical Latents and Masked AutoregRessive Modeling

HALO introduces a hyperspherical VAE to constrain continuous latent representations to a fixed‑radius shell, stabilizing numerical fluctuations. It then employs a masked autoregressive model that balances parallel decoding with temporal correlation learning, reducing inference steps and improving stability. Experiments show HALO achieves state‑of‑the‑art generation performance with significantly better inference efficiency compared to existing baselines.

By Chunyi Hou, Xiangfei Qiu, Hanyin Cheng, Yutong Li, Bin Yang
Hugging Face Trending Papers
Aug 13

V-RAE: Rethinking Video Latent Spaces for Generation

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.