arXiv Machine Learning
Sep 24

Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders

The paper introduces LLMAE, a technique that transforms a pretrained decoder-only language model into a continuous text autoencoder by inserting a fixed-length latent bottleneck into its internal activations. Using a 270M Gemma 3 model with structured attention masks, LoRA adaptation, and KL regularization, LLMAE achieves near-perfect reconstruction of text sequences up to 1024 tokens. The authors further show that the resulting latent representation can be leveraged to train a latent text diffusion model for detailed image captioning, demonstrating downstream utility.

By Arkanath Pathak, Unnat Jain, Alexander C. Berg