The paper introduces Successive Capacity Growth (SCG), a method for adaptively expanding Vision Transformer encoders in Joint-Embedding Predictive Architectures (JEPAs). SCG starts with a minimal encoder and incrementally increases width or depth based on a task‑agnostic test‑and‑verify mechanism, while a Sketched Isotropic Gaussian Regularizer (SIGReg) keeps learned semantic dimensions independent. Experiments on multi‑object dynamics and 2D navigation tasks show that SCG achieves up to 20.3% better prediction loss than fixed small baselines and 23% better than fixed large models, with far greater parameter efficiency and no false‑positive expansions.
By Frederik Berenz
LoopVAE introduces a recurrent depth architecture that reuses a scale‑ and loop‑conditioned core across different spatial scales while keeping resolution‑changing transitions separate. The four‑block core applies 28 block operations per encoder or decoder, enabling a 29M‑parameter convolutional model to achieve 0.28 rFID and 32.54 dB PSNR on ImageNet‑256 with roughly 65% fewer parameters than comparable VAEs. Experiments with both convolutional and Transformer operators, as well as ablations on parameter sharing, demonstrate competitive image quality metrics and reveal how targeted loop interventions and truncation affect reconstruction quality and computational trade‑offs.
By Zhiying Lu
arXiv:2609.31620v1 Announce Type: new
Abstract: Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual r...
By Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui, Dawei Liu, Runhao Li, Jingcheng Ni, Chen Wei, Randall Balestriero, Yue Wang
arXiv:2605. 18324v2 Announce Type: replace-cross Abstract: Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders.
By Jaskirat Singh, Boyang Zheng, Zongze Wu, Richard Zhang, Eli Shechtman, Saining Xie
The paper introduces the Superposed Latent Autoencoder (SLAE), a method that stores multiple wide latent representations together by superposing them into a single memory tensor using learned codes and randomized keys. SLAE eliminates the need for tight dimensional bottlenecks, achieving up to 56% lower reconstruction error on datasets such as CIFAR-10/100 and SVHN while maintaining the same storage budget. The approach also boosts downstream classification performance by up to 16.79 percentage points, demonstrating that wide representations can be effectively compressed through structured interference rather than dimensional reduction.
By Quanling Zhao, Jiaying Yang, Tianqi Zhang, Ziyang Hao, Fatemeh Asgarinejad, Flavio Ponzina, Tajana Rosing
arXiv:2606. 16112v1 Announce Type: cross Abstract: Residual architectures are ubiquitous in deep learning, but they suffer from a subtle structural limitation: the norm of the residual stream can grow rapidly with depth.
By Tom\'as Figliolia, Beren Millidge
arXiv:2602. 07697v3 Announce Type: replace-cross Abstract: Predictive coding (PC) is a biologically plausible alternative to standard backpropagation (BP) that minimises an energy function with respect to network activities before updating weights.
By Francesco Innocenti, El Mehdi Achour, Rafal Bogacz
arXiv:2606. 14990v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are standard tools for mechanistic interpretability, but current SAE families are constrained by fixed encoder nonlinearities such as ReLU, JumpReLU, and TopK.
By Naiyu Yin, Yue Yu
The paper introduces Successive Capacity Growth (SCG), a method that starts with a minimal Vision Transformer encoder and incrementally expands its width or depth based on a task‑agnostic test‑and‑verify mechanism. SCG uses function‑preserving expansion and a Sketched Isotropic Gaussian Regularizer (SIGReg) to ensure independent semantic dimensions and prevent collapse. Experiments on multi‑object dynamics and 2D navigation tasks show that SCG achieves significant prediction loss reductions while being far more parameter‑efficient than fixed large models, with no false‑positive expansions and exact function preservation.
arXiv:2609.05730v1 Announce Type: cross
Abstract: Contrastive Language-Image Pretraining (CLIP) is a building block of many machine learning applications. Scaling laws have guided resource allocation...
By Samir Char, Carles Domingo-Enrich, Randall Balestriero
arXiv:2603. 04198v2 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) are widely used to extract human-interpretable features from neural network activations, but their learned features can vary substantially across random seeds and training choices.
By Piotr Jedryszek, Oliver M. Crook
arXiv:2607. 08605v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept.
By Weiduo Liao, Yunqiao Yang, Ying Wei