arXiv AI

HiRAE: Hierarchical Representation Autoencoding with Residual Budgets

arXiv Computer Vision
3d ago

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

arXiv:2609.31620v1 Announce Type: new Abstract: Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual r...

By Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui, Dawei Liu, Runhao Li, Jingcheng Ni, Chen Wei, Randall Balestriero, Yue Wang
Hugging Face Trending Papers
Jun 9

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing semantically rich latent spaces for image generation. However, their reconstruction quality often remains suboptimal, largely because deep VFM representations do not preserve sufficient fine-grained visual detail.

arXiv Machine Learning
Sep 24

Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders

The paper introduces LLMAE, a technique that transforms a pretrained decoder-only language model into a continuous text autoencoder by inserting a fixed-length latent bottleneck into its internal activations. Using a 270M Gemma 3 model with structured attention masks, LoRA adaptation, and KL regularization, LLMAE achieves near-perfect reconstruction of text sequences up to 1024 tokens. The authors further show that the resulting latent representation can be leveraged to train a latent text diffusion model for detailed image captioning, demonstrating downstream utility.

By Arkanath Pathak, Unnat Jain, Alexander C. Berg
arXiv AI
Jul 2

UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios

arXiv:2511. 18050v1 Announce Type: cross Abstract: Diffusion transformers have recently delivered strong text-to-image generation around 1K resolution, but we show that extending them to native 4K across diverse aspect ratios exposes a tightly coupled failure mode spanning positional encoding, VAE compression, and optimization.

By Tian Ye, Song Fei, Lei Zhu