Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data. This fade is captured by an envelope $f(\ell)$.
arXiv:2606. 29519v1 Announce Type: new Abstract: Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data.
By Lorenzo Livi
arXiv:2606. 30789v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard tool for improving the reasoning ability of large language models, yet its training dynamics are still described empirically: reward trajectories are fit with low-parameter functional forms whose constants carry no mechanistic meaning, and hyperparameter choices remain a matter of trial and error.
By Rajat Ghosh, Datta Nimmaturi, Aryan Singhal, Vaishnavi Bhargava, Henry Wong, Johnu George, Debojyoti Dutta
The paper introduces a world model that learns to predict the evolution of physical systems while respecting key physical principles. By hard‑coding a general structure—generating dynamics from the gradient of a learned energy via a fixed reversible operator and imposing constraints on energy, dissipation, and interventions—the model achieves second‑law compatible dissipation, accurate responses to parameter changes, long‑term stability, and robustness to disturbances. Experiments on an electromagnetic cavity, a particle‑in‑cell grid, and shallow‑water fluid demonstrate that the model can recover accurate constitutive functions, distinguish conserving from dissipating regimes, and transfer learned physics to unseen conditions, outperforming unconstrained models.
By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling
arXiv:2507.10383v5 Announce Type: replace-cross
Abstract: Recurrent neural networks are canonical models of biological memory. In these models, memories are represented by distributed patterns of neu...
By Uri Cohen, M\'at\'e Lengyel
arXiv:2606. 18080v1 Announce Type: new Abstract: Gradient descent in deep learning may operate at the edge of stability (EoS), a regime in which the largest eigenvalue of the loss Hessian hovers near the stability threshold $2/\eta$, where $\eta$ is the learning rate.
By Pierre Marion
arXiv:2607. 09714v1 Announce Type: new Abstract: The Feedback-Coupled Memory Systems (FCMS) architecture formalizes closed-loop coordination through four abstract operators, two of which - the agent update operator $f_i$ and the environmental update operator $\Psi$ - are left axiomatically undefined in the original framework.
By Stefano Grassi
arXiv:2608. 00097v1 Announce Type: cross Abstract: Physical learning rules such as equilibrium propagation (EP), coupled learning (CL), and adjoint coupled learning (AL) train resistive networks through local measurements.
By Bijaya Dangol
The paper demonstrates that learned simulators can fail in two distinct ways when conditions change: long‑horizon drift due to accumulated errors and incorrect responses to interventions on physical parameters. By adding a symplectic integrator to preserve conservative dynamics, rollouts remain stable for up to 100× the training horizon, while encoding physical coupling via explicit linear factorization allows the model to generalize to unseen signs of that coupling. The study shows that stability and counterfactual generalization arise from separate structural choices, enabling designers to impose each property independently.
By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling
The paper derives an exact discrete‑time law that captures how learning‑rate schedules and weight decay interact in scale‑invariant neural networks, showing that a single scalar quantity governs the effective step size. It demonstrates that the balance point between contraction and expansion is intrinsically unstable, leading to recurrent dynamics when using constant learning rates with weight decay. The authors extend this analysis to various optimizers and datasets, confirming the law’s precision and showing that performance peaks sharply at the predicted boundary.
By Hasan Amin, Wei-Kai Chang, Rajiv Khanna
arXiv:2609.16827v1 Announce Type: new
Abstract: High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit exceptional storage capabilities and robustness. Previous empirica...
By Akira Tamamori
The paper investigates how two independent inductive biases—one from the circuit’s invariance under conductance rescaling and one from the learning rule’s conservation of a mass quantity—affect what a physical learning system remembers. By separating these effects, the authors show that when every element is trainable, the initialization scale has negligible influence on the learned function, whereas a single untrainable element can cause the function to shift significantly with initialization. They further demonstrate that the conservation law does not protect memory but instead influences solution quality, with adjoint coupled learning (AL) generally performing worse than equilibrium propagation (EP) and coupled learning (CL) in small circuits.
whyItMatters":"The study clarifies that only the circuit’s structural bias, not the rule’s conservation property, determines memory retention in physical learning systems."
By Bijaya Dangol