arXiv:2609.31620v1 Announce Type: new
Abstract: Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual r...
By Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui, Dawei Liu, Runhao Li, Jingcheng Ni, Chen Wei, Randall Balestriero, Yue Wang
arXiv:2609.37775v1 Announce Type: cross
Abstract: Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. M...
By Xuanyu Zhu, Yan Bai, Yang Shi, Yihang Lou, Yuanxing Zhang, Tengfei Liu, Jing Jin, Yuan Zhou
The paper introduces LLMAE, a technique that transforms a pretrained decoder-only language model into a continuous text autoencoder by inserting a fixed-length latent bottleneck into its internal activations. Using a 270M Gemma 3 model with structured attention masks, LoRA adaptation, and KL regularization, LLMAE achieves near-perfect reconstruction of text sequences up to 1024 tokens. The authors further show that the resulting latent representation can be leveraged to train a latent text diffusion model for detailed image captioning, demonstrating downstream utility.
By Arkanath Pathak, Unnat Jain, Alexander C. Berg
arXiv:2410. 07299v3 Announce Type: replace-cross Abstract: We introduce OTIS, an open time series encoder that yields high-quality time series features for downstream deployment on any system, including resource-constrained wearables and industrial sensors.
By \"Ozg\"un Turgut, Philip M\"uller, Martin J. Menten, Daniel Rueckert
arXiv:2607. 08605v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept.
By Weiduo Liao, Yunqiao Yang, Ying Wei
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.