arXiv:2602. 07488v3 Announce Type: replace-cross Abstract: Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset.
By Francesco Cagnetta, Allan Ravent\'os, Surya Ganguli, Matthieu Wyart
arXiv:2609.25797v1 Announce Type: new
Abstract: We reply to the comments by M. Sienicki and K. Sienicki (arXiv:2512.07881) and by K. Sienicki (arXiv:2601.06104) on our work on quantum-mechanical stat...
By Massimiliano Sassoli de Bianchi, Roberto Leporini
arXiv:2602. 03685v2 Announce Type: replace-cross Abstract: Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origin remains debatable.
By Yizhou Liu, Ziming Liu, Cengiz Pehlevan, Jeff Gore
arXiv:2606. 06238v1 Announce Type: new Abstract: We propose a statistical-field framework for text generated by large language models (LLMs), treating token embeddings as continuous spin variables on a one-dimensional chain.
By Huajian Ruan, Jinyang Li, Xingyu Guo, Lingxiao Wang
arXiv:2606. 24504v1 Announce Type: new Abstract: We discuss reasons why the scaling exponents of current Large Language Models (LLMs) applications are indicating an unsustainable regime in terms of energy resources.
By Sauro Succi, Peter V. Coveney, Alex Hansen
arXiv:2606. 25008v1 Announce Type: new Abstract: Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute.
By Yizhou Liu, Jeff Gore
arXiv:2510. 06945v2 Announce Type: replace Abstract: Motivated by the growing interest in quantum machine learning, in particular quantum neural networks (QNNs), we study how recently introduced evaluation metrics based on the Fisher information matrix (FIM) are effective for predicting their training and prediction performance.
By Lorenzo Pastori, Veronika Eyring, Mierk Schwabe
arXiv:2511.00763v3 Announce Type: replace
Abstract: We investigate the performance of large language models (LLMs) on repetitive deterministic prediction tasks and study how the sequence accuracy rat...
By Wanda Hou, Leon Zhou, Hong-Ye Hu, Yubei Chen, Yi-Zhuang You, Xiao-Liang Qi
arXiv:2602. 18364v2 Announce Type: replace-cross Abstract: Maximum likelihood prediction (MLP) is a core task at the heart of modern large language models.
By Sreejith Sreekumar, Nir Weinberger
arXiv:2605. 25344v2 Announce Type: replace-cross Abstract: Dense linear maps carry much of the parameter and computational burden of modern neural networks, yet their dense form leaves the organization of learned couplings implicit.
By Ying Lu, Peng-Fei Zhou, Qi-Xuan Fang, Pan Zhang, Shi-Ju Ran, Gang Su
arXiv:2608. 14691v1 Announce Type: new Abstract: Sequence models are conventionally distinguished by their backbone, the mechanism that routes information across positions, such as attention or recurrence.
By Ahmed Nebli, Hadi Saadatdoorabi, Christopher Keibel, Kevin Yam
arXiv:2602. 06699v2 Announce Type: replace-cross Abstract: We propose a variational quantum implementation of self-attention (QSA)-the core operation in transformers and large language models-which predicts future elements of a sequence by forming overlap-weighted combinations of past data.
By Alessio Pecilli, Matteo Rosati