Semantic Chunking and the Entropy of Natural Language
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 11371v1 Announce Type: cross Abstract: Spoken language, whether produced by humans or large language models (LLM), unfolds over time with varying semantic content.
The paper offers a unified probabilistic framework for large language models, describing them as probability measures over token sequences defined by autoregressive conditional distributions. Training is cast as maximum‑likelihood estimation solved via stochastic gradient methods, while generation is treated as sequential simulation of the resulting stochastic process. It also explores how the asymmetry of the Kullback–Leibler divergence relates to hallucination and the distinction between plausibility and truth, and extends the perspective to diffusion models that generate data by simulating a reverse‑time stochastic process.
arXiv:2406. 05335v3 Announce Type: replace-cross Abstract: Generation of text and speech in natural languages can be modeled as a stochastic process.
The paper proposes an information-weighted cross‑entropy loss that rescales token contributions using TF‑IDF statistics, thereby emphasizing semantically informative tokens and down‑weighting ubiquitous ones. Experiments on five decoder‑only language models (1.1B–13B parameters) show consistent reductions in memorized substring length while maintaining perplexity and downstream performance. The method is architecture‑agnostic, adds less than 3% computational overhead, and can be integrated into existing training pipelines.
arXiv:2605.27268v2 Announce Type: replace-cross Abstract: Modern Large Language Models (LLMs) are often criticized for producing repetitive and homogeneous text, despite possessing vast latent vocabu...
arXiv:2605. 05103v3 Announce Type: replace-cross Abstract: We introduce the \textbf{Concept Field} of a text corpus: a local drift field with pointwise uncertainty, estimated in sentence-embedding space from the deltas between consecutive sentences.