arXiv Computer Vision By Xianghong Fang, Wenjie Shu, Tongda Xu, Wenlong Mou, Dehan Kong, Tim G. J. Rudner

Bridging Reconstruction and Generation: A Latent Distribution Perspective on Evaluation and Improvement

Read the original on arXiv Computer Vision →

The paper investigates why reconstruction quality in latent generative models does not always predict generative performance, attributing the issue to a mismatch between encoder-induced and generation-time latent distributions. It introduces Generation‑Aware Reconstruction (GAR), a method that perturbs encoder latents with noise and denoises them through the generative model before decoding, creating a continuous trajectory that reveals how the decoder behaves across latent spaces. The resulting GAR‑FID metric correlates strongly with generation FID, and using intermediate GAR latents for decoder adaptation consistently improves generative quality across different model scales.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 3

Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies

The paper introduces Latent Drift, a generative forecasting framework that predicts slow-evolving neurodegenerative disease progression by learning changes in a compressed semantic representation rather than full-resolution anatomy. It addresses two failure modes—identity collapse and continuous interpolation trap—by removing pixel-level identity from the prediction target and applying Finite Scalar Quantization to suppress high-frequency nuisance fluctuations. Experiments on longitudinal 3D brain MRI demonstrate that Latent Drift outperforms diffusion and autoregressive transformer baselines in both generative fidelity and clinically relevant metrics.

By Yuxiang Feng, Juncheng Wang, Chao Xu, Wenlong Hou, Huihan Wang, Yijie Qian, Yang Liu, Baigui Sun, Yong Liu, Shujun Wang
arXiv AI
Aug 12

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".

By Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit