Geometry-Preserving Encoder/Decoder in Latent Generative Models
arXiv:2501. 09876v3 Announce Type: replace-cross Abstract: Generative modeling aims to generate new data samples that resemble a given dataset.
arXiv:2606. 15553v1 Announce Type: cross Abstract: Representation Autoencoders (RAEs) have improved diffusion and flow models by semantically richer latent space owing to the strongly label-wise clustered DINO features in the pretrained encoders.
arXiv:2501. 09876v3 Announce Type: replace-cross Abstract: Generative modeling aims to generate new data samples that resemble a given dataset.
arXiv:2607. 29180v1 Announce Type: cross Abstract: Text-to-motion generation must produce motions that are semantically correct, temporally coherent, and physically plausible.
arXiv:2410. 10137v5 Announce Type: replace Abstract: We develop Riemannian approaches to variational autoencoders (VAEs) for PDE-type ambient data with regularizing geometric latent dynamics, which we refer to as VAE-DLM, or VAEs with dynamical latent manifolds.
The paper investigates how fine‑tuning pretrained visual encoders for faithful image reconstruction affects diffusion models that operate in the resulting latent space. It finds that such fine‑tuning reduces the effective dimensionality of the latent representation, causing standard velocity‑prediction flow‑matching to fit noise outside the low‑dimensional signal manifold and making optimization inefficient. Consequently, the authors propose using a clean‑data ($oldsymbol{x}_{0}$) parameterization, which focuses learning on the signal manifold and consistently improves text‑to‑image generation across multiple strong‑reconstruction encoders.
arXiv:2609.40037v1 Announce Type: new Abstract: Few-step autoregressive video generation enables efficient streaming synthesis, but errors introduced in early temporal blocks are reused as context an...
arXiv:2605. 12183v2 Announce Type: replace Abstract: Drifting Models have emerged as a new paradigm for one-step generative modeling, achieving strong image quality without iterative inference.
arXiv:2609.36348v1 Announce Type: cross Abstract: Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas th...
arXiv:2606. 07036v1 Announce Type: cross Abstract: Synthetic histopathology image generation addresses critical challenges in computational pathology, including patient privacy and the growing need for large-scale training data for foundation models.
arXiv:2605. 18324v2 Announce Type: replace-cross Abstract: Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders.
Unified image restoration (UIR) aims to recover high-quality (HQ) content from low-quality (LQ) images with different degradations using a single model. Most recent methods adapt large pretrained text-to-image (T2I) latent diffusion models for their strong capacity and generative priors.
arXiv:2608. 01298v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks.
arXiv:2605. 20708v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) have become a de facto backbone of modern visual generation, and nearly every major axis of their design -- tokenization, attention, conditioning, objectives, and latent autoencoders -- has been extensively revisited.