arXiv Machine Learning

A Hamiltonian-Inspired Local-Operator Ansatz for Slimming Large Language Models

arXiv:2605. 25344v2 Announce Type: replace-cross Abstract: Dense linear maps carry much of the parameter and computational burden of modern neural networks, yet their dense form leaves the organization of learned couplings implicit.

arXiv Machine Learning
Aug 31

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

The paper investigates the often-overlooked scale vectors in large language models, showing that despite their tiny size they are crucial for pre‑training performance. The authors provide theoretical insights that scale vectors mainly aid optimization rather than expressivity, and they analyze how weight decay affects different normalization layers. Building on these findings, they propose lightweight improvements—branch‑specific heterogeneity, better placement, and magnitude‑direction reparameterization—that consistently reduce loss across a range of model sizes and training settings.

By Mingze Wang, Shuchen Zhu, Yuxin Fang, Binghui Li, Kai Shen, Shu Zhong
arXiv Machine Learning
Aug 19

Center-Manifold Reduction of Learning at Bifurcations: Interference and Rich Learning in Recurrent Neural Networks

The paper investigates how gradient descent behaves near codimension‑one bifurcations in recurrent neural networks by analyzing the global empirical Neural Tangent Kernel (GeNTK). Under local center‑manifold conditions, the parameter‑to‑state Jacobian is approximated by a low‑rank normal‑form operator, causing the GeNTK and Fisher information matrix to become strongly amplified and anisotropic, concentrating on a rank‑one or rank‑two channel depending on the bifurcation type. Experiments on high‑dimensional RNNs confirm that this low‑rank concentration coincides with abrupt loss changes, subtask interference, and aligns with changes in memory dynamics in a 15‑task LeakyRNN.

By James Hazelden, Eric Shea-Brown
arXiv Machine Learning
Jul 24

Cautious optimism for deep parameterized quantum circuits

arXiv:2607. 21409v1 Announce Type: cross Abstract: A central challenge in quantum machine learning is understanding the scaling behavior of parameterized quantum circuits (PQCs).

By Marie Kempkes, Elies Gil-Fuster, Carlos Bravo-Prieto, Aroosa Ijaz, Alissa Wilms, Jens Eisert, Evert van Nieuwenburg, Vedran Dunjko