arXiv Machine Learning By Kai Nakaishi, Yoshihiko Nishikawa, Koji Hukushima

Phase transition in large language models and the criticality of natural languages

Read the original on arXiv Machine Learning →

arXiv:2406. 05335v3 Announce Type: replace-cross Abstract: Generation of text and speech in natural languages can be modeled as a stochastic process.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Aug 18

Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation

The paper argues that language operates with two parameters: amplitude, which measures how often words co‑occur, and phase, a signed relational factor that determines how co‑activated meanings combine and can reverse a meaning’s contribution. Unlike amplitude, phase is not captured by standard word embeddings or transformer attention weights and is indexed to individuals and dyadic interactions. The authors propose six empirical predictions to test phase’s role and suggest that future language models should incorporate agent‑indexed, phase‑bearing semantic states.

arXiv Computation and Language
Sep 18

Foundations of Stochastic Lexical Calculus: Semantic Descent and Random Dynamics on Probability Simplices

The paper introduces a framework called stochastic lexical calculus that determines when probabilities produced by large language models can be used to represent sequential states in scientific systems. It defines typed measurable transformations of contextual language, constructs a minimal closed representation, and provides necessary and sufficient conditions for unique semantic updates. The authors prove bounds on irreducible nonclosure and accumulated error, and show that under average contraction an external random recursion on a probability simplex is stable and unique. Empirical tests on frozen experiments demonstrate that raw prompt-conditioned probabilities fail an invariance gate, but after prompt-specific calibration a common three-state representation satisfies stability gates and covers 28 of 30 eight-step paths, achieving 0.933 coverage at a nominal 0.90 level.

By Matthew F Dixon
arXiv AI
Jul 10

How Do I Know What to Say Next? Barenholtz's Autogenerative Theory as an Enrichment of Harrisean Integrationism

arXiv:2607. 07891v1 Announce Type: cross Abstract: Roy Harris's Integrationist linguistics offers a compelling critique of the referentialist tradition embedded deep at the heart of computational approaches to language, arguing that language is not a code that maps onto a pre-given world but a situated, bipartite activity oriented toward prospective joint action.

By J. Mark Bishop, Stephen J. Cowley