arXiv AI By Zhaotian Gu, Jie Su, Weiwei Wang, Chang Liu, Tianyi Qian, Dahui Wang

Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory

Read the original on arXiv AI →

The paper introduces the Recurrent Divisive Normalization Network (RDNN), a minimal model that incorporates divisive normalization—a common neural computation—to stabilize continuous working memory representations. Dynamical systems analysis shows that this biophysical constraint enables the network to converge to robust, high‑fidelity slow manifolds, while gradient dynamics during Backpropagation Through Time reveal an activity‑dependent local scaling that compresses the network’s effective rank into a low‑dimensional subspace. Ablation studies confirm that divisive normalization, rather than subtractive inhibition, is essential for preventing manifold shattering under time‑varying inputs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 19

Center-Manifold Reduction of Learning at Bifurcations: Interference and Rich Learning in Recurrent Neural Networks

The paper investigates how gradient descent behaves near codimension‑one bifurcations in recurrent neural networks by analyzing the global empirical Neural Tangent Kernel (GeNTK). Under local center‑manifold conditions, the parameter‑to‑state Jacobian is approximated by a low‑rank normal‑form operator, causing the GeNTK and Fisher information matrix to become strongly amplified and anisotropic, concentrating on a rank‑one or rank‑two channel depending on the bifurcation type. Experiments on high‑dimensional RNNs confirm that this low‑rank concentration coincides with abrupt loss changes, subtask interference, and aligns with changes in memory dynamics in a 15‑task LeakyRNN.

By James Hazelden, Eric Shea-Brown
arXiv AI
Sep 25

ELiSe: Efficient Learning of Sequences in Structured Recurrent Networks

The paper introduces ELiSe, a model that leverages cortical network scaffolds and dendritic compartments to learn complex non‑Markovian spatio‑temporal patterns using only local, always‑on, phase‑free synaptic plasticity. It demonstrates the model’s ability to acquire and replay intricate sequences, exemplified by a birdsong learning mock‑up, and shows robustness to external disturbances and flexibility in parameter settings.

By Laura Kriener, Kristin V\"olk, Ben von H\"unerbein, Federico Benitez, Walter Senn, Mihai A. Petrovici
Hugging Face Trending Papers
Sep 8

When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay

The paper investigates how normalization makes neural networks scale‑invariant, creating a feedback loop between learning‑rate schedules and weight decay that controls the effective step size of the optimizer. It derives an exact discrete‑time law showing that a single scalar quantity captures all schedule and decay effects, with norm growth providing a self‑quenching counter‑force that defines a sharp boundary between contraction‑ and expansion‑dominated regimes. Through exact analysis of a normalized regression model and experiments on MLPs, CNNs, GPT‑2, and various datasets, the authors demonstrate that constant learning rates with weight decay are intrinsically unstable, leading to recurrent dynamics, and that adaptive optimizers exhibit weaker stabilization under normalization. "whyItMatters":"The study provides a precise, actionable rule for controlling training dynamics and schedule design in modern deep learning by isolating a single governing quantity for scale‑invariant optimization."