arXiv AI

Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory

The paper introduces the Recurrent Divisive Normalization Network (RDNN), a minimal model that incorporates divisive normalization—a common neural computation—to stabilize continuous working memory representations. Dynamical systems analysis shows that this biophysical constraint enables the network to converge to robust, high‑fidelity slow manifolds, while gradient dynamics during Backpropagation Through Time reveal an activity‑dependent local scaling that compresses the network’s effective rank into a low‑dimensional subspace. Ablation studies confirm that divisive normalization, rather than subtractive inhibition, is essential for preventing manifold shattering under time‑varying inputs.

arXiv Machine Learning
Aug 19

Center-Manifold Reduction of Learning at Bifurcations: Interference and Rich Learning in Recurrent Neural Networks

The paper investigates how gradient descent behaves near codimension‑one bifurcations in recurrent neural networks by analyzing the global empirical Neural Tangent Kernel (GeNTK). Under local center‑manifold conditions, the parameter‑to‑state Jacobian is approximated by a low‑rank normal‑form operator, causing the GeNTK and Fisher information matrix to become strongly amplified and anisotropic, concentrating on a rank‑one or rank‑two channel depending on the bifurcation type. Experiments on high‑dimensional RNNs confirm that this low‑rank concentration coincides with abrupt loss changes, subtask interference, and aligns with changes in memory dynamics in a 15‑task LeakyRNN.

By James Hazelden, Eric Shea-Brown
arXiv AI
Sep 25

ELiSe: Efficient Learning of Sequences in Structured Recurrent Networks

The paper introduces ELiSe, a model that leverages cortical network scaffolds and dendritic compartments to learn complex non‑Markovian spatio‑temporal patterns using only local, always‑on, phase‑free synaptic plasticity. It demonstrates the model’s ability to acquire and replay intricate sequences, exemplified by a birdsong learning mock‑up, and shows robustness to external disturbances and flexibility in parameter settings.

By Laura Kriener, Kristin V\"olk, Ben von H\"unerbein, Federico Benitez, Walter Senn, Mihai A. Petrovici
Hugging Face Trending Papers
Sep 8

When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay

The paper investigates how normalization makes neural networks scale‑invariant, creating a feedback loop between learning‑rate schedules and weight decay that controls the effective step size of the optimizer. It derives an exact discrete‑time law showing that a single scalar quantity captures all schedule and decay effects, with norm growth providing a self‑quenching counter‑force that defines a sharp boundary between contraction‑ and expansion‑dominated regimes. Through exact analysis of a normalized regression model and experiments on MLPs, CNNs, GPT‑2, and various datasets, the authors demonstrate that constant learning rates with weight decay are intrinsically unstable, leading to recurrent dynamics, and that adaptive optimizers exhibit weaker stabilization under normalization. "whyItMatters":"The study provides a precise, actionable rule for controlling training dynamics and schedule design in modern deep learning by isolating a single governing quantity for scale‑invariant optimization."

arXiv Machine Learning
Sep 10

When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay

The paper derives an exact discrete‑time law that captures how learning‑rate schedules and weight decay interact in scale‑invariant neural networks, showing that a single scalar quantity governs the effective step size. It demonstrates that the balance point between contraction and expansion is intrinsically unstable, leading to recurrent dynamics when using constant learning rates with weight decay. The authors extend this analysis to various optimizers and datasets, confirming the law’s precision and showing that performance peaks sharply at the predicted boundary.

By Hasan Amin, Wei-Kai Chang, Rajiv Khanna
arXiv Machine Learning
Aug 6

Contrastive Diffusion Alignment: Learning Structured Latents for Controllable Generation

arXiv:2510. 14190v3 Announce Type: replace Abstract: Diffusion models excel at generation, but their latent spaces are high dimensional and not explicitly organized for interpretation or control.

By Ruchi Sandilya, Sumaira Perez, Charles Lynch, Lindsay Victoria, Benjamin Zebley, Derrick Matthew Buchanan, Mahendra T. Bhati, Nolan Williams, Timothy J. Spellman, Faith M. Gunning, Conor Liston, Logan Grosenick
Hugging Face Trending Papers
Jul 15

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip connections and normalization, long understood as controlling magnitude, as mechanisms for preserving gradient rank across depth, since the very matrix multiplications and nonlinear activations that make the network expressive also reduce the rank.

arXiv Machine Learning
Jun 4

Drift-Diffusion Matching: Embedding dynamics in latent manifolds of asymmetric neural networks

arXiv:2602. 14885v2 Announce Type: replace-cross Abstract: Recurrent neural networks (RNNs) provide a theoretical framework for understanding computation in biological neural circuits, yet classical results, such as Hopfield's model of associative memory, rely on symmetric connectivity that restricts network dynamics to gradient-like flows.

By Ram\'on Nartallo-Kaluarachchi, Renaud Lambiotte, Alain Goriely