arXiv Machine Learning By Giorgio Racca, Michal Valko, Amartya Sanyal

Language Generation with Replay: A Learning-Theoretic View of Model Collapse

Read the original on arXiv Machine Learning →

arXiv:2603. 11784v2 Announce Type: replace Abstract: As scaling laws push the training of frontier large language models (LLMs) toward ever-growing data requirements, training pipelines are approaching a regime where much of the publicly available online text may be consumed.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 30

Generating in the Limit with Infinitely Many Hallucinations

arXiv:2606. 28354v1 Announce Type: cross Abstract: The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown target language, and a learner tasked with identifying that language.

By Irene Strauss, Alexandra Butoi, Ryan Cotterell
arXiv AI
Jun 9

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

arXiv:2602. 08222v2 Announce Type: replace Abstract: As post-training optimization becomes central to improving large language models, we observe a persistent saturation bottleneck: once models grow highly confident, further training yields diminishing returns.

By Zehao Chen, Gongxun Li, Tianxiang Ai, Zixuan Huang, Xiaodong Liu, Yifei Li, Wang Zhou, Fuzhen Zhuang, Xianglong Liu, Jianxin Li, Deqing Wang, Yikun Ban