arXiv:2609.38149v1 Announce Type: new
Abstract: Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for informa...
By Dor Tirosh, Ido Amos, Mor Geva
arXiv:2603. 07523v3 Announce Type: replace Abstract: Transferring knowledge by fine-tuning large-scale pre-trained networks has become a standard paradigm for downstream tasks, yet the knowledge of a pre-trained model is tightly coupled with monolithic architecture, which restricts flexible reuse across models of varying scales.
By Jianlu Shen, Fu Feng, Yucheng Xie, Jiaqi Lv, Xin Geng
arXiv:2607. 07743v1 Announce Type: cross Abstract: Self-organization is an emergent property of life, driven by the collective behavior of individual components acting on local information.
By Meet Barot, Daniel Berenberg, Sina Khajehabdollahi
The paper proposes a method for task adaptation that eliminates the need for gradient computation during adaptation. Using a Neural Cellular Automaton, the authors train recurrent dynamics and memory read/write operations via backpropagation, then fix the slow model parameters. Online adaptation is achieved solely through local memory updates driven by prediction errors, enabling significant performance gains on new classification tasks with a single support set pass.
By Krsto Prorokovi\'c
The paper introduces Regularized Latent Dynamics Prediction (RLDP), a method that adds orthogonality regularization to self‑supervised next‑state prediction in latent space. RLDP maintains feature diversity, matching or surpassing complex representation learning approaches for zero‑shot reinforcement learning. It also performs robustly in low‑coverage data settings where prior methods fail.
By Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang, Scott Niekum, Martha White
arXiv:2606. 04048v1 Announce Type: cross Abstract: Training and scaling Large Language Models demand enormous computational resources, motivating both efficient sub-quadratic architectures and principled hyperparameter tuning methods.
By Yifeng Liu, Quanquan Gu
arXiv:2609.37631v1 Announce Type: new
Abstract: Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work sh...
By Zachary Shinnick, Christian Intern\`o, Hemanth Saratchandran, Anton van den Hengel, Damien Teney
arXiv:2602. 20062v2 Announce Type: replace Abstract: Pretraining and fine-tuning are central stages in modern machine learning systems.
By Nicolas Anguita, Francesco Locatello, Andrew M. Saxe, Marco Mondelli, Flavia Mancini, Samuel Lippl, Clementine Domine
arXiv:2606. 24969v1 Announce Type: new Abstract: While the quadratic sequence-length bottleneck of transformers has fueled a resurgence in recurrent models, effectively capturing complex dynamics requires architectures that balance efficient training with highly expressive latent states.
By Klaus Schertler, Xiomara Runge, Andrea Ceni, David Kappel, Claudio Gallicchio
The paper introduces ELiSe, a model that leverages cortical network scaffolds and dendritic compartments to learn complex non‑Markovian spatio‑temporal patterns using only local, always‑on, phase‑free synaptic plasticity. It demonstrates the model’s ability to acquire and replay intricate sequences, exemplified by a birdsong learning mock‑up, and shows robustness to external disturbances and flexibility in parameter settings.
By Laura Kriener, Kristin V\"olk, Ben von H\"unerbein, Federico Benitez, Walter Senn, Mihai A. Petrovici
AYLA is a loss reparameterization framework that applies a sigmoid‑controlled power‑law transformation to the empirical loss, dynamically adjusting gradient magnitudes without changing stationary points or optimal solutions. By reshaping optimization trajectories, AYLA accelerates descent in flat or saddle‑dominated regions and stabilizes late‑stage training, leading to improved feature recovery in two‑layer tanh networks on synthetic Gaussian data. Experiments show enhanced weight alignment, neuron similarity, activation correlation, and richer internal representations, while mitigating rank collapse and promoting a transition from lazy to active feature‑learning regimes.
By Behnam Gheshlaghi, Shahin Atakishiyev
arXiv:2606. 17120v1 Announce Type: new Abstract: Deep neural networks (DNNs) exhibit first order phase transitions under variations of the L2 regularization strength, with each transition marking the onset of a new learnable feature.
By Ibrahim Talha Ersoy, Karoline Wiesner