The paper investigates why large language models (LLMs) fail to maintain continuous mixtures of token embeddings—used in latent-state reasoning—to preserve multiple reasoning paths. Through theory and experiments, it identifies three failure sources: transformer geometry distortion, amplification or contraction dynamics from softmax and autoregressive feedback, and the need for context-dependent corrections that scale with mixture size. Empirical results confirm the predicted transition between contraction and amplification and show pretrained models largely fall on the amplifying side.
By Ali Backour
arXiv:2606. 07559v1 Announce Type: cross Abstract: Fine-tuning a language model on contexts whose correct completion has a near-synonym competitor often fails silently.
By Vaibhav Prakash, Jayasri Dontabhaktuni
arXiv:2606. 07559v2 Announce Type: replace-cross Abstract: Fine-tuning a language model often fails silently when its correct completion must outrank a near-synonym competitor.
By Vaibhav Prakash, Jayasri Dontabhaktuni
arXiv:2609.07474v2 Announce Type: replace
Abstract: Language models compute over tokens: language is their input, their output, and increasingly their internal representation. Whether language should...
By Peng Xie, Amr Alanwar
arXiv:2607. 16741v1 Announce Type: new Abstract: B\"urger et al.
By Francesco Karim Vicidomini
arXiv:2607. 18305v1 Announce Type: cross Abstract: Some limits on what language models know are not gaps in data coverage but structural properties of learning from text.
By Priyansh Srivastava, Romit Chatterjee
arXiv:2609.01170v1 Announce Type: new
Abstract: Large language models exhibit a modular internal organization that mirrors well-studied functional networks of the human brain, but how this organizati...
By Guangqi Li, Yongxin Li
The paper investigates how recursive contamination—retraining language models on their own generated text—affects output diversity across 13 publicly released checkpoints. Using a fixed contamination protocol over five generations, the authors find a wide spread in 4‑gram diversity (0.187 to 0.940), indicating that some models collapse into repetitive fragments while others remain largely unaffected. The study shows that a model’s susceptibility to collapse is an intrinsic property of the checkpoint, not predicted by parameter scale or static indicators, and that simple interventions such as tightening top‑p sampling can significantly slow or halt collapse.
By Yangze Liu, Zhongyi Han
arXiv:2608. 02830v1 Announce Type: cross Abstract: Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied.
By Mohammad Rostami
arXiv:2609.14384v1 Announce Type: new
Abstract: What must a neural system be capable of to implement language? Current research annotates stimuli with linguistic variables and tests which electrodes,...
By Elliot Murphy
Neural Collapse predicts that balanced one-hot classification pushes model representations to be equally far from each other; a symmetric configuration that depends only on the output label and ignores any semantic similarity in the inputs. This creates a puzzle: next-token prediction language models are trained predominantly (as context length increases) with one-hot labels: the same context is very unlikely to appear twice in training with different labels.
arXiv:2606. 24752v1 Announce Type: new Abstract: The loss of plasticity - the ability of a network to learn new information after having already learned older information - is a fundamental challenge in creating artificial neural networks capable of continual learning.
By J. Fernando Hernandez-Garcia, Tom\'as Figliolia, Beren Millidge