arXiv Computation and Language
2d ago

Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling

The paper proposes a new architecture for masked language modeling that replaces the Transformer attention mechanism with a stack of low‑rank bottleneck autoencoders. Each autoencoder mixes information locally, across the full sequence, and across attention heads, compressing and reconstructing inputs without training‑dependent width. An iterative refinement process at masked positions pulls embeddings toward a weighted neighbor average and then projects them back onto the learned manifold, achieving comparable performance to BERT with roughly 1.9× fewer FLOPs and matching BERT on rare‑token performance through a frequency‑aware training schedule.

By Narges Mokhtari, Farzan Haddadi, Ebrahim Rezaii