Hugging Face Trending Papers

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining

Midway through an ordinary pretraining run, a small language model learns the pronoun-gender rule: cued with a girl's name ("Sue cried because"), it resolves the next pronoun to she, generalizing to held-out probes (0. 94 by step 925).

arXiv Machine Learning
Sep 11

A Fragility Spectrum for Recursive Language-Model Training

The paper investigates how recursive contamination—retraining language models on their own generated text—affects output diversity across 13 publicly released checkpoints. Using a fixed contamination protocol over five generations, the authors find a wide spread in 4‑gram diversity (0.187 to 0.940), indicating that some models collapse into repetitive fragments while others remain largely unaffected. The study shows that a model’s susceptibility to collapse is an intrinsic property of the checkpoint, not predicted by parameter scale or static indicators, and that simple interventions such as tightening top‑p sampling can significantly slow or halt collapse.

By Yangze Liu, Zhongyi Han
arXiv Computation and Language
Sep 25

What a Cross-Model Fixed-Point Census Can and Cannot Arbitrate About Repetition

The paper investigates neural text degeneration by measuring the fixed‑point structure of short‑window argmax maps across 17 pretrained models, using 96 random two‑token starts without prompts. It finds a stable four‑way classification that varies across model families and scales, with some models funneling to a single endpoint token while others do not, and shows that this behavior is not solely determined by training data or corpus frequency. The study demonstrates that repetition phenomena are not uniformly explained by either training data or network architecture alone, highlighting the complexity of neural text generation dynamics.

By Nicol\'as Vera Z\'u\~niga
arXiv Machine Learning
Sep 17

Capability Emergence Can Be Forecast: Per-Seed, In Advance, With Calibrated Intervals, Certified False Alarms, and a Blind Pre-Registered Gate

The paper demonstrates that emergent capabilities in machine learning models can be forecasted with lead time, calibrated uncertainty, and controlled false‑alarm rates. Using per‑seed analysis on transformers, the authors show that the formation time of a previous‑token head predicts the emergence of an induction head with Spearman ρ = 0.977 and a median lead of 975 training steps. Conformal intervals, blind pre‑registered tests, and a multiplicative rule relating anchor and event times further validate the predictive framework across multiple model families and configurations.

By Gunner Levi Howe
arXiv AI
Aug 28

What the "Spotless" Mind Remembers: How Knowledge Entanglement Shapes What Leaks After Unlearning in LLMs

The paper investigates how the structural entanglement of facts within a large language model’s knowledge base influences whether those facts leak after unlearning. Using two unlearning algorithms (WHP and GA+KL) across fictional and real-world datasets and multiple model sizes, the authors find that highly entangled facts are more likely to be recalled before unlearning, but the relationship changes—WHP weakens it while GA+KL reverses it. By directly manipulating entanglement scores and observing corresponding recall changes, they demonstrate a causal link and develop a predictive tool to audit prompts for potential leakage.

By Aakriti Shah, Yifan Hu, Thai Le
arXiv Machine Learning
Aug 4

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit

arXiv:2608. 02302v1 Announce Type: cross Abstract: Long-horizon coding-agent trajectories are poorly matched to the credit units available to train on: a single action has no stable value, an episode label merges productive exploration with abandoned directions, and a fixed window cuts where the logging mechanics fall.

By Jingxi Wei