Hugging Face Trending Papers

Can Scale Save Us From Plasticity Loss in Large Language Models?

The loss of plasticity - the ability of a network to learn new information after having already learned older information - is a fundamental challenge in creating artificial neural networks capable of continual learning. Although this phenomenon has been known for decades, it has mostly been studied in older, relatively small architectures and rarely in natural-language domains.

arXiv AI
Jun 24

Can Scale Save Us From Plasticity Loss in Large Language Models?

arXiv:2606. 24752v1 Announce Type: new Abstract: The loss of plasticity - the ability of a network to learn new information after having already learned older information - is a fundamental challenge in creating artificial neural networks capable of continual learning.

By J. Fernando Hernandez-Garcia, Tom\'as Figliolia, Beren Millidge
arXiv AI
Sep 10

Limits of Reliability and Scaling in Language Models

The paper argues that large language models cannot achieve perfect reliability for any task, even with unlimited scale. It establishes that each generative task has an inherent reliability ceiling set by how much output uncertainty can be resolved from observable context, with a resolvable part that can be improved by more context and a subjective part tied to task ambiguity. The authors derive a scaling law showing that performance is limited by the scarcer resource—either training data or model capacity—and explain how this law explains phenomena such as retrieval augmentation and catastrophic forgetting.

By Subhabrata Majumdar
arXiv Machine Learning
Jun 25

Internal Data Repetition Destroys Language Models

arXiv:2606. 24998v1 Announce Type: new Abstract: Language models are running out of high-quality training data, and even aggressively deduplicated corpora retain some amount of repetition.

By Jessica Chudnovsky, Joshua Kazdan, Noam Levi, Rylan Schaeffer, Yegor Denisov-Blanch, Bo He, Mehmet Donmez, Sanmi Koyejo, David Donoho
arXiv Computation and Language
Sep 16

Persistent Recurrent Memory Between Transformer Layers - Improves Language Model Generalization

The paper proposes a lightweight recurrent memory module inserted between the lower and upper halves of a 6‑layer decoder‑only transformer. This module, which uses cross‑attention to observe hidden states, a GRU to update a persistent state, and gated addition to modulate subsequent layers, adds only 3.7% more parameters. It reduces evaluation loss by 28.5% and narrows the generalization gap, with ablations showing the benefit comes solely from the memory topology rather than auxiliary losses.

By Eduardo Novaes Hering
arXiv Machine Learning
Aug 31

Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

The paper investigates how learning rate and batch size scale when pretraining dense large language models on English‑prevalent corpora, examining both jointly optimal and marginal evolutions across model capacity and data size. It explores the benefits of a Warmup‑Stable‑Decay learning‑rate schedule, assessing whether optimal hyperparameters transfer between stable and decay phases, and evaluates loss scaling forms that capture interactions between model capacity and dataset size. The study provides a baseline scaling procedure and releases the full set of pretraining runs for future OpenEuroLLM development.

By Niccol\`o Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan, J\"org Franke, Sampo Pyysalo, Jenia Jitsev, Aaron Klein
arXiv AI
Sep 1

On the Plasticity Collapse in Continual Machine Unlearning

The paper investigates continual machine unlearning, where models must forget data over time. It identifies a fundamental issue called plasticity collapse, where successive unlearning requests cause geometric constraints that saturate parameter space, leading to two failure modes: forward failure (reduced forgetting quality) and backward failure (re‑memorization). Experiments across architectures and datasets confirm that plasticity collapse is a pervasive problem in continual unlearning.

By Yingdan Shi, Xiang Xu, Kaize Ding, Alfred O. Hero, Ren Wang
arXiv AI
Aug 20

Forgetting, plasticity, and co-observation: a third facet of continual learning

The paper argues that catastrophic forgetting and loss of plasticity alone cannot explain why naive sequential training underperforms offline joint training. It introduces data co-observation as a third factor, showing that observing training data together consistently improves performance across supervised and self-supervised settings. The study also reinterprets common continual learning methods, suggesting that memory replay’s success stems from restoring co-observation benefits rather than merely mitigating forgetting.

By Timm Hess, Abhishek Jha, Gido M. van de Ven, Tinne Tuytelaars
arXiv Machine Learning
Sep 3

The Dynamics of Continuous Mixture Collapse in Language Models

The paper investigates why large language models (LLMs) fail to maintain continuous mixtures of token embeddings—used in latent-state reasoning—to preserve multiple reasoning paths. Through theory and experiments, it identifies three failure sources: transformer geometry distortion, amplification or contraction dynamics from softmax and autoregressive feedback, and the need for context-dependent corrections that scale with mixture size. Empirical results confirm the predicted transition between contraction and amplification and show pretrained models largely fall on the amplifying side.

By Ali Backour