arXiv AI

Semantic Chunking and the Entropy of Natural Language

arXiv Machine Learning
Sep 23

The Probabilistic Structure of Large Language Models

The paper offers a unified probabilistic framework for large language models, describing them as probability measures over token sequences defined by autoregressive conditional distributions. Training is cast as maximum‑likelihood estimation solved via stochastic gradient methods, while generation is treated as sequential simulation of the resulting stochastic process. It also explores how the asymmetry of the Kullback–Leibler divergence relates to hallucination and the distinction between plausibility and truth, and extends the perspective to diffusion models that generate data by simulating a reverse‑time stochastic process.

By Adnan Aboulala\^a
arXiv Machine Learning
Sep 11

Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss

The paper proposes an information-weighted cross‑entropy loss that rescales token contributions using TF‑IDF statistics, thereby emphasizing semantically informative tokens and down‑weighting ubiquitous ones. Experiments on five decoder‑only language models (1.1B–13B parameters) show consistent reductions in memorized substring length while maintaining perplexity and downstream performance. The method is architecture‑agnostic, adds less than 3% computational overhead, and can be integrated into existing training pipelines.

By Zhijian Li, Stefan Larson, Kevin Leach
Hugging Face Trending Papers
Jun 25

Structure Before Collapse: Transient semantic geometry in next-token prediction

Neural Collapse predicts that balanced one-hot classification pushes model representations to be equally far from each other; a symmetric configuration that depends only on the output label and ignores any semantic similarity in the inputs. This creates a puzzle: next-token prediction language models are trained predominantly (as context length increases) with one-hot labels: the same context is very unlikely to appear twice in training with different labels.