Foundations of Large Language Models
Foundations of Large Language Models is a book that focuses on core concepts of large language models rather than exhaustive coverage of the latest technologies. It is organized into six chapters covering pre‑training, generative models, prompting, alignment, inference, and reasoning. The book targets college students, professionals, and practitioners in NLP and related fields, serving as a reference for anyone interested in large language models.
Related stories
A Survey on Diffusion Language Models
arXiv:2508. 10875v3 Announce Type: replace-cross Abstract: Diffusion Language Models (DLMs) are rapidly emerging as a powerful and promising alternative to the dominant autoregressive (AR) paradigm.
Efficient training of language models to fill in the middle
Cosmopedia: how to create large-scale synthetic data for pre-training Large Language Models
Large Language Models: A Mathematical Formulation
The article presents a mathematical framework for large language models (LLMs), detailing how text sequences are encoded into tokens, how next‑token prediction architectures are defined, and how these models are trained and deployed for tasks such as summarization, recommendation, software writing, and quantitative problem solving. It emphasizes that the framework relies on basic concepts from information theory, probability, and optimization, yet captures the complex algorithmic structure responsible for LLMs’ empirical successes. The authors argue that this formalism enables the study of accuracy, efficiency, and robustness, and points toward new methodological developments.
All Entities are Not Created Equal: Examining the Long Tail for Ultra-Fine Entity Typing
arXiv:2410.17355v4 Announce Type: replace Abstract: Due to their capacity to acquire world knowledge from large corpora, pre-trained language models (PLMs) are extensively used in ultra-fine entity t...
The State Of LLMs 2025: Progress, Problems, and Predictions
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
Learning to Translate from Soft to Hard LLM Prompts
arXiv:2605. 27642v2 Announce Type: replace-cross Abstract: Soft prompting, also known as continuous prompting, is a parameter-efficient method for tuning LLMs to specific tasks.
Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC)
The paper introduces LOGIC (Logit‑Space Integration for Contextual Biasing), a new framework that injects contextual entity information directly into the decoding layer of Speech Large Language Models, bypassing the limitations of prompt‑based methods. LOGIC operates with constant‑time complexity regardless of the size of the entity list, and experiments with the Phi‑4‑MM model across 11 multilingual locales show an average 9% relative reduction in Entity WER while adding only a 0.30% increase in False Alarm Rate.
Mimir: Large-scale Multilingual Concept Modeling
arXiv:2605.25263v2 Announce Type: replace-cross Abstract: Current language modeling approaches are built around tokens. Text corpora are split into tokens, and models are trained by performing comput...
Structured Inference with Large Language Gibbs
arXiv:2606. 19264v1 Announce Type: new Abstract: The knowledge encoded in large language models (LLMs) can serve as a substrate for structured reasoning over variables describing a complex world, but accessing this knowledge in a probabilistically coherent manner poses a difficult inference problem.
Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
arXiv:2605. 17932v2 Announce Type: replace-cross Abstract: Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus mainly on autoregressive architectures.
