arXiv AI

Neural Cellular Automata Learn General Features in their Hidden Channels

Neural Cellular Automata (NCAs) are shown to learn general, scale‑invariant topological primitives in their hidden channels, which can be transferred from a teacher to a student model for few‑shot learning. The study introduces a transfer‑learning mechanism that injects pretrained hidden states into a student, improving early optimization and outperforming recurrent and feed‑forward baselines on MNIST benchmarks with only ~9,800 parameters. Mechanistic analysis reveals that hidden channels decouple feature extraction from classification, converging to mutually orthogonal states that absorb morphological complexity.

arXiv Machine Learning
Jun 4

Breaking the Scale Barrier: One-Shot Knowledge Transfer via Frequency Transform

arXiv:2603. 07523v3 Announce Type: replace Abstract: Transferring knowledge by fine-tuning large-scale pre-trained networks has become a standard paradigm for downstream tasks, yet the knowledge of a pre-trained model is tightly coupled with monolithic architecture, which restricts flexible reuse across models of varying scales.

By Jianlu Shen, Fu Feng, Yucheng Xie, Jiaqi Lv, Xin Geng
arXiv AI
Jul 10

Architecture Generalization with MetaNCA

arXiv:2607. 07743v1 Announce Type: cross Abstract: Self-organization is an emergent property of life, driven by the collective behavior of individual components acting on local information.

By Meet Barot, Daniel Berenberg, Sina Khajehabdollahi
arXiv Machine Learning
Sep 25

Online Task Adaptation via Self-Organisation

The paper proposes a method for task adaptation that eliminates the need for gradient computation during adaptation. Using a Neural Cellular Automaton, the authors train recurrent dynamics and memory read/write operations via backpropagation, then fix the slow model parameters. Online adaptation is achieved solely through local memory updates driven by prediction errors, enabling significant performance gains on new classification tasks with a single support set pass.

By Krsto Prorokovi\'c
arXiv Machine Learning
Aug 27

Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models

The paper introduces Regularized Latent Dynamics Prediction (RLDP), a method that adds orthogonality regularization to self‑supervised next‑state prediction in latent space. RLDP maintains feature diversity, matching or surpassing complex representation learning approaches for zero‑shot reinforcement learning. It also performs robustly in low‑coverage data settings where prior methods fail.

By Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang, Scott Niekum, Martha White
arXiv Machine Learning
Jun 25

Frequency Domain Reservoir Computing

arXiv:2606. 24969v1 Announce Type: new Abstract: While the quadratic sequence-length bottleneck of transformers has fueled a resurgence in recurrent models, effectively capturing complex dynamics requires architectures that balance efficient training with highly expressive latent states.

By Klaus Schertler, Xiomara Runge, Andrea Ceni, David Kappel, Claudio Gallicchio
arXiv AI
Sep 25

ELiSe: Efficient Learning of Sequences in Structured Recurrent Networks

The paper introduces ELiSe, a model that leverages cortical network scaffolds and dendritic compartments to learn complex non‑Markovian spatio‑temporal patterns using only local, always‑on, phase‑free synaptic plasticity. It demonstrates the model’s ability to acquire and replay intricate sequences, exemplified by a birdsong learning mock‑up, and shows robustness to external disturbances and flexibility in parameter settings.

By Laura Kriener, Kristin V\"olk, Ben von H\"unerbein, Federico Benitez, Walter Senn, Mihai A. Petrovici
arXiv Machine Learning
5d ago

AYLA: Architecting a loss landscape in shallow neural networks to accelerate feature recovery

AYLA is a loss reparameterization framework that applies a sigmoid‑controlled power‑law transformation to the empirical loss, dynamically adjusting gradient magnitudes without changing stationary points or optimal solutions. By reshaping optimization trajectories, AYLA accelerates descent in flat or saddle‑dominated regions and stabilizes late‑stage training, leading to improved feature recovery in two‑layer tanh networks on synthetic Gaussian data. Experiments show enhanced weight alignment, neuron similarity, activation correlation, and richer internal representations, while mitigating rank collapse and promoting a transition from lazy to active feature‑learning regimes.

By Behnam Gheshlaghi, Shahin Atakishiyev