arXiv Statistics ML By Jorge Medina Moreira, Lorenzo Bardone, Lenka Zdeborov\'a

The hidden advantage of mask resampling: a theory of masked autoencoders

Read the original on arXiv Statistics ML →

arXiv:2610. 01578v1 Announce Type: new Abstract: Why can masked prediction learn useful representations that unmasked reconstruction misses?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Statistics ML.

arXiv Computer Vision
6d ago

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

arXiv:2609.31620v1 Announce Type: new Abstract: Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual r...

By Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui, Dawei Liu, Runhao Li, Jingcheng Ni, Chen Wei, Randall Balestriero, Yue Wang
arXiv Computation and Language
6d ago

Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling

The paper proposes a new architecture for masked language modeling that replaces the Transformer attention mechanism with a stack of low‑rank bottleneck autoencoders. Each autoencoder mixes information locally, across the full sequence, and across attention heads, compressing and reconstructing inputs without training‑dependent width. An iterative refinement process at masked positions pulls embeddings toward a weighted neighbor average and then projects them back onto the learned manifold, achieving comparable performance to BERT with roughly 1.9× fewer FLOPs and matching BERT on rare‑token performance through a frequency‑aware training schedule.

By Narges Mokhtari, Farzan Haddadi, Ebrahim Rezaii