arXiv:2608. 12925v1 Announce Type: new Abstract: Momentum-based optimizers are widely used in modern deep learning, yet the relations among momentum recursion, update geometry, and acceleration remain only partially understood.
By Zhixin Ren, Yau Lyu, Congrong Li, Liping Zhang, Shengbo Eben Li
arXiv:2602. 10204v2 Announce Type: replace Abstract: We introduce MVN-Grad (Momentum on Variance-Normalized Gradients), an Adam-style optimizer that improves stability and performance by combining two complementary ideas: variance-based normalization and momentum applied after normalization.
By Francisco Patitucci, Aryan Mokhtari
The paper introduces Gradient‑Momentum Coupling (GMC), a method that quantifies learning progress by measuring how strongly a sample influences changes in the parameter space, using the normalized absolute product of its gradient and the momentum of previous gradients. GMC filters out noise by accumulating consistent directions of change while canceling random fluctuations, leading to a more uniform prioritization across tasks with varying noise levels and better ranking of learnable tasks by improvement speed. Experiments on MiniGrid MultiRoom tasks show that replacing prediction error with GMC in the Intrinsic Curiosity Module restores exploration capabilities that were lost to unpredictable observations.
By Samuel Blad, Martin L\"angkvist, Amy Loutfi
arXiv:2606. 08783v1 Announce Type: cross Abstract: Orthogonalized momentum updates, as used in Muon-style optimizers, have recently shown strong empirical stability in large-scale deep learning.
By Ganzhao Yuan
arXiv:2609.07534v1 Announce Type: cross
Abstract: Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sustained c...
By Yuhan Wang, Yurou Chen, Hongye Jiang, Wenzhao Lian
arXiv:2507. 14056v3 Announce Type: replace-cross Abstract: Recent work in continual learning has highlighted the stability gap -- a temporary performance drop on previously learned tasks when new ones are introduced.
By Alejandro Rodriguez-Garcia, Anindya Ghosh, Srikanth Ramaswamy
arXiv:2609.07534v2 Announce Type: replace-cross
Abstract: Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sus...
By Yuhan Wang, Yurou Chen, Hongye Jiang, Wenzhao Lian
arXiv:2609.36738v1 Announce Type: new
Abstract: Optimizer momentum is usually stored as a parameter-sized moving average of past gradients, which makes history costly and fixes each past signal in th...
By Yuchen Li, Zongqi Fan, Nguyen H. Tran, Ken-Tye Yong
arXiv:2609.17042v1 Announce Type: new
Abstract: Learning flexible motor primitives is a hallmark of skilled motor control. Recent neuroscience theory proposes that motor primitives may be implemented...
By Sreejan Kumar, Marcelo Mattar, Lea Duncker
arXiv:2607. 17257v1 Announce Type: cross Abstract: Diffusion policies have shown strong potential for robotic imitation learning, and recent extensions incorporate additional modalities to improve manipulation performance.
By Zihao He, Hongjie Fang, Shirun Tang, Cewu Lu, Haoshu Fang
arXiv:2607. 15880v1 Announce Type: cross Abstract: Imitation Learning aims to learn skills from extensive observations and demonstrations for robots, so it suffers from data scarcity and environment generalization.
By Zhenduo Shang, Xiyao Liu, Bohan Li, Xudong Wang, Teng Ren, Lianqing Liu, Zhi Han
arXiv:2604. 18933v2 Announce Type: replace-cross Abstract: Robotic manipulation tasks exhibit varying memory requirements, ranging from Markovian tasks that require no memory to non-Markovian tasks that demand in-context memorization of historical information within a single trial or in-context adaptation based on the outcomes of multiple past trials.
By Yihuai Gao, Jeff Jinyun Liu, Shuang Li, Shuran Song