arXiv Machine Learning

Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders

arXiv:2602. 10099v2 Announce Type: replace Abstract: Leveraging representation encoders for generative modeling offers a path for efficient, high-fidelity synthesis.

Hugging Face Trending Papers
Jul 27

RODR: Riemannian Orthogonally Decoupled Regularization for Disentangled Manifold Representation

Point cloud denoising is essentially a geometric recovery task that aims to reconstruct the intrinsic structure of a smooth 2D Riemannian manifold embedded in R^3 from noisy, discrete ambient-space samples. Despite the remarkable progress of modern manifold-aware encoders and generative transport models in geometric representation learning, a fundamental objective-geometry mismatch remains underexplored.

arXiv Machine Learning
Sep 24

On the Diffusibility of High-Dimensional Latents

The paper investigates how fine‑tuning pretrained visual encoders for faithful image reconstruction affects diffusion models that operate in the resulting latent space. It finds that such fine‑tuning reduces the effective dimensionality of the latent representation, causing standard velocity‑prediction flow‑matching to fit noise outside the low‑dimensional signal manifold and making optimization inefficient. Consequently, the authors propose using a clean‑data ($oldsymbol{x}_{0}$) parameterization, which focuses learning on the signal manifold and consistently improves text‑to‑image generation across multiple strong‑reconstruction encoders.

By Chao Feng, Zhiyang Xu, Bowei Chen, Yuanjun Xiong, Xiyao Wang, Jui-Hsien Wang, Richard Zhang, Zhe Lin, Andrew Owens, Yijun Li
Hugging Face Trending Papers
Aug 5

Intrinsic-Hybrid Latent Diffusion Models for Generative Modeling on Unknown Manifolds

We introduce the Intrinsic Hybrid Latent Diffusion Model (ILDM), a generative framework that integrates probabilistic dimensionality reduction with geometry-aware diffusion on unknown manifolds. While diffusion models (DMs) have achieved state-of-the-art results in high-dimensional data synthesis, they rely on large training datasets and ignore intrinsic geometric structure.

arXiv Computer Vision
Sep 4

Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers

The paper introduces a novel compression framework for image-to-shape Diffusion Transformers (DiTs) that significantly reduces model size while preserving geometric fidelity. By exploiting the non-uniform importance of 3D DiT layers, the authors combine structured pruning, adaptive quantization, and targeted fine‑tuning into a vitality‑guided approach. The method achieves up to a 66% reduction in model size across state‑of‑the‑art image‑to‑3D models without compromising synthesis quality, offering a plug‑and‑play solution for efficient 3D shape generation.

By Jaeah Lee, Hyunjin Kim, Jaewoong Cho, Gihyun Kwon