arXiv Machine Learning

Large Language Models: A Mathematical Formulation

The article presents a mathematical framework for large language models (LLMs), detailing how text sequences are encoded into tokens, how next‑token prediction architectures are defined, and how these models are trained and deployed for tasks such as summarization, recommendation, software writing, and quantitative problem solving. It emphasizes that the framework relies on basic concepts from information theory, probability, and optimization, yet captures the complex algorithmic structure responsible for LLMs’ empirical successes. The authors argue that this formalism enables the study of accuracy, efficiency, and robustness, and points toward new methodological developments.

arXiv Machine Learning
Sep 23

The Probabilistic Structure of Large Language Models

The paper offers a unified probabilistic framework for large language models, describing them as probability measures over token sequences defined by autoregressive conditional distributions. Training is cast as maximum‑likelihood estimation solved via stochastic gradient methods, while generation is treated as sequential simulation of the resulting stochastic process. It also explores how the asymmetry of the Kullback–Leibler divergence relates to hallucination and the distinction between plausibility and truth, and extends the perspective to diffusion models that generate data by simulating a reverse‑time stochastic process.

By Adnan Aboulala\^a
arXiv AI
Sep 25

Foundations of Large Language Models

Foundations of Large Language Models is a book that focuses on core concepts of large language models rather than exhaustive coverage of the latest technologies. It is organized into six chapters covering pre‑training, generative models, prompting, alignment, inference, and reasoning. The book targets college students, professionals, and practitioners in NLP and related fields, serving as a reference for anyone interested in large language models.

By Tong Xiao, Jingbo Zhu
arXiv Machine Learning
Aug 4

Just on Time: Token-Level Early Stopping for Diffusion Language Models

arXiv:2602. 11133v2 Announce Type: replace Abstract: Diffusion language models generate text through iterative refinement, a process that is often computationally inefficient because many tokens reach stability long before the final denoising step.

By Zakhar Kohut, Severyn Shykula, Mykola Vysotskyi, Serhii Dmytryshyn, Dmytro Khamula, Michal Zakrzewski, Damian Rynczak, Jacek Ma{\l}ecki, Taras Rumezhak, Volodymyr Karpiv
arXiv Machine Learning
Jun 4

Transmuting prompts into weights

arXiv:2510. 08734v3 Announce Type: replace Abstract: A growing body of research has demonstrated that the behavior of large language models can be effectively controlled at inference time by directly modifying their internal states, either through vector additions to their activations or through updates to their weight matrices.

By Hanna Mazzawi, Benoit Dherin, Michael Munn, Adrian Goldwaser, Michael Wunder, Javier Gonzalvo
arXiv Computation and Language
Sep 16

TIAO: Token Importance-Aware Policy Optimization for Text Summarization

The paper introduces Token Importance-Aware Policy Optimization (TIAO), a reinforcement learning approach that improves text summarization by weighting token importance based on token dependency. TIAO reweights a trajectory’s advantage according to the overall dependencies of core tokens, addressing the limitation of previous methods that treat all tokens equally. Experiments demonstrate that a 7B foundation model enhanced with TIAO achieves performance comparable to GPT‑4 and GPT‑5‑nano on real‑world datasets.

By Qixiu Li, Chenlong Bao, Xiang Zhu, Xiaoyong Li, Ruixin Cao, Shukai Chen, Zhenxiong Zhou