arXiv:2609.37775v1 Announce Type: cross
Abstract: Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. M...
By Xuanyu Zhu, Yan Bai, Yang Shi, Yihang Lou, Yuanxing Zhang, Tengfei Liu, Jing Jin, Yuan Zhou
arXiv:2607.14088v2 Announce Type: replace
Abstract: Video generation models typically rely on 3D-VAEs trained for pixel-level reconstruction, whose latent spaces may underrepresent semantic structure...
By Zhihao Xie, Junfeng Wu, Xinting Hu, Junchao Huang, Li Jiang
arXiv:2609.36756v1 Announce Type: cross
Abstract: One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressiv...
By Jiawei Zhang, Shuhao Liu, Rong Huang, Yuancheng Li, Zhihui Li, Xiaojun Chang, Changlin Li
One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation qual...
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.
arXiv:2602.20731v2 Announce Type: replace-cross
Abstract: Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or...
By Aram Davtyan, Yusuf Sahin, Yasaman Haghighi, Sebastian Stapf, Pablo Acuaviva, Alexandre Alahi, Paolo Favaro