Midway through an ordinary pretraining run, a small language model learns the pronoun-gender rule: cued with a girl's name ("Sue cried because"), it resolves the next pronoun to she, generalizing to held-out probes (0. 94 by step 925).
The paper investigates how recursive contamination—retraining language models on their own generated text—affects output diversity across 13 publicly released checkpoints. Using a fixed contamination protocol over five generations, the authors find a wide spread in 4‑gram diversity (0.187 to 0.940), indicating that some models collapse into repetitive fragments while others remain largely unaffected. The study shows that a model’s susceptibility to collapse is an intrinsic property of the checkpoint, not predicted by parameter scale or static indicators, and that simple interventions such as tightening top‑p sampling can significantly slow or halt collapse.
By Yangze Liu, Zhongyi Han
The paper investigates neural text degeneration by measuring the fixed‑point structure of short‑window argmax maps across 17 pretrained models, using 96 random two‑token starts without prompts. It finds a stable four‑way classification that varies across model families and scales, with some models funneling to a single endpoint token while others do not, and shows that this behavior is not solely determined by training data or corpus frequency. The study demonstrates that repetition phenomena are not uniformly explained by either training data or network architecture alone, highlighting the complexity of neural text generation dynamics.
By Nicol\'as Vera Z\'u\~niga
arXiv:2609.11149v3 Announce Type: replace-cross
Abstract: How fast does a language model degrade when trained on its own outputs? Theory traces it to gradually accumulating errors, while experiments...
By Yangze Liu, Zhongyi Han
arXiv:2609.39243v1 Announce Type: new
Abstract: Causal interventions such as activation patching and distributed alignment search (DAS) are the main tool for making mechanistic claims about neural ne...
By Beiming Liu, Minjie Chen
arXiv:2607. 25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves.
By Cen Lu, Yung-Chen Tang, Andrea Cavallaro
arXiv:2609. 31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts.
By Nicol\'as Vera Z\'u\~niga
The paper demonstrates that emergent capabilities in machine learning models can be forecasted with lead time, calibrated uncertainty, and controlled false‑alarm rates. Using per‑seed analysis on transformers, the authors show that the formation time of a previous‑token head predicts the emergence of an induction head with Spearman ρ = 0.977 and a median lead of 975 training steps. Conformal intervals, blind pre‑registered tests, and a multiplicative rule relating anchor and event times further validate the predictive framework across multiple model families and configurations.
By Gunner Levi Howe
arXiv:2510.17021v2 Announce Type: replace-cross
Abstract: Large language model (LLM) unlearning is a key approach for removing undesired data, knowledge, or behaviors from pretrained models while ret...
By Bingqi Shang, Yiwei Chen, Yihua Zhang, Bingquan Shen, Sijia Liu
arXiv:2602. 02470v2 Announce Type: replace Abstract: Autoregressive large language models (LLMs) have achieved remarkable success in many complex tasks, yet they can still fail in very simple logical reasoning such as the "reversal curse" -- when trained on forward knowledge data of the form "$A \rightarrow B$" (e.
By Xutao Ma, Yixiao Huang, Hanlin Zhu, Somayeh Sojoudi
arXiv:2608. 02302v1 Announce Type: cross Abstract: Long-horizon coding-agent trajectories are poorly matched to the credit units available to train on: a single action has no stable value, an episode label merges productive exploration with abandoned directions, and a fixed window cuts where the logging mechanics fall.
By Jingxi Wei
arXiv:2608. 13063v1 Announce Type: new Abstract: Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies.
By Sam Mao