arXiv:2609.38149v1 Announce Type: new
Abstract: Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for informa...
By Dor Tirosh, Ido Amos, Mor Geva
arXiv:2603. 07523v3 Announce Type: replace Abstract: Transferring knowledge by fine-tuning large-scale pre-trained networks has become a standard paradigm for downstream tasks, yet the knowledge of a pre-trained model is tightly coupled with monolithic architecture, which restricts flexible reuse across models of varying scales.
By Jianlu Shen, Fu Feng, Yucheng Xie, Jiaqi Lv, Xin Geng
arXiv:2607. 07743v1 Announce Type: cross Abstract: Self-organization is an emergent property of life, driven by the collective behavior of individual components acting on local information.
By Meet Barot, Daniel Berenberg, Sina Khajehabdollahi
The paper proposes a method for task adaptation that eliminates the need for gradient computation during adaptation. Using a Neural Cellular Automaton, the authors train recurrent dynamics and memory read/write operations via backpropagation, then fix the slow model parameters. Online adaptation is achieved solely through local memory updates driven by prediction errors, enabling significant performance gains on new classification tasks with a single support set pass.
By Krsto Prorokovi\'c
The paper introduces Regularized Latent Dynamics Prediction (RLDP), a method that adds orthogonality regularization to self‑supervised next‑state prediction in latent space. RLDP maintains feature diversity, matching or surpassing complex representation learning approaches for zero‑shot reinforcement learning. It also performs robustly in low‑coverage data settings where prior methods fail.
By Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang, Scott Niekum, Martha White
arXiv:2606. 04048v1 Announce Type: cross Abstract: Training and scaling Large Language Models demand enormous computational resources, motivating both efficient sub-quadratic architectures and principled hyperparameter tuning methods.
By Yifeng Liu, Quanquan Gu