arXiv:2512. 22088v3 Announce Type: replace-cross Abstract: The scaling law, a cornerstone of Large Language Model (LLM) development, predicts improvements in model performance with increasing computational resources.
By Chiwun Yang
arXiv:2609.27581v1 Announce Type: new
Abstract: Step Law gives power-law formulas for the optimal peak learning rate eta* and batch size B* when pre-training language models. It was calibrated on mod...
By Egor Romanyukov, Timofey Novikov, Timur Shokarov, Elizaveta Zorkina, Anastasia Palienko, Stepan Dergachev
arXiv:2607. 22757v1 Announce Type: cross Abstract: We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention, and the training objective.
By T. Shaska
The paper presents empirical scaling laws for autoregressive language models, linking prediction loss to model size, data size, and compute, and investigates their theoretical basis using a teacher–student linear RNN framework. In this tractable setting, a stable latent linear RNN generates trajectories while a sketched linear recurrent student is trained via full‑batch WSD gradient descent on next‑token prediction. The study derives explicit approximation, optimization, and statistical scaling laws that depend on the sketch dimension, number of trajectories, and trajectory length, revealing how different power‑law exponents for innovation and initialization covariances affect the rates and crossovers between regimes.
By Ziyan Chen, Zhongzhu Zhou, Peilin Liu, Ding-Xuan Zhou
Informed Masking (IM) is a new technique for aligning Diffusion Large Language Models (dLLMs) with Reinforcement Learning (RL). It identifies a systematic upstream/downstream token structure in dLLM rollouts and shows that masking downstream tokens creates better subproblems for likelihood estimation. When integrated into three state‑of‑the‑art dLLM RL methods on LLaDA‑8B‑Instruct, IM yields up to 2.01%, 8.68%, and 5.77% relative average gains on math and planning benchmarks while improving training stability.
By Xiaoyi Yu, Enver Sangineto, Pei Fu, Fiorenzo Parascandolo, Wenhui Tan, Ruikang Zhang, Rita Cucchiara, Ruihua Song, Jian Luan
arXiv:2505. 24275v4 Announce Type: replace Abstract: We propose GradPower, a lightweight gradient-transformation technique for accelerating language model pre-training.
By Jinbo Wang, Mingze Wang, Jiaqi Zhang, Wei Wang, Peng Pei, Xunliang Cai, Weinan E, Lei Wu
arXiv:2609.37745v1 Announce Type: new
Abstract: Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a $1/3$ exponent...
By Hyunseok Lee, Mihir Basil, Yizhou Liu, Jeff Gore
arXiv:2607. 10848v1 Announce Type: new Abstract: Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize off-policy updates.
By Xiangxin Zhou, Jiarui Yao, Penghui Qi, Bowen Ping, Jiaqi Tang, Haonan Wang, Tianyu Pang
arXiv:2609.37717v1 Announce Type: new
Abstract: Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state thro...
By Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu State
arXiv:2609.37535v1 Announce Type: new
Abstract: In the softmax output layer, a rare token receives a small positive logit gradient on most steps and a much larger negative gradient on the few steps w...
By Sangsidhya Kar
The paper argues that Large Language Models (LLMs) do not function as Solomonoff induction estimators because their training objectives—cross‑entropy, negative log‑likelihood, and next‑token prediction—optimize fit to a supplied conditional distribution rather than a program‑weighted universal mixture. It further contends that additional computation alone does not transform these models into optimal predictors without external hyper‑parameter or architectural changes. The authors suggest that neurosymbolic machine learning, exemplified by models such as Fable and Astra, represents a shift toward symbolic model synthesis, moving beyond purely statistical LLMs.
By Hector Zenil, Abicumaran Uthamacumaran, Luan Ozelim
arXiv:2605. 25085v2 Announce Type: replace-cross Abstract: We study the rate-distortion limits of online KV cache compression in autoregressive language models, formulating it as sequential Wyner-Ziv source coding on the filtration induced by the model, with the next-step query as decoder side information.
By Munsik Kim