Epistemic diversity across language models mitigates knowledge collapse
arXiv:2512. 15011v3 Announce Type: replace Abstract: Artificial intelligence (AI) increasingly generates the very content used to train future AI systems.
arXiv:2605. 04127v2 Announce Type: replace Abstract: Model collapse, the degradation in performance that arises when generative models are trained on the outputs of prior models, is an increasing concern as artificially generated content proliferates.
arXiv:2512. 15011v3 Announce Type: replace Abstract: Artificial intelligence (AI) increasingly generates the very content used to train future AI systems.
arXiv:2608.21366v1 Announce Type: new Abstract: Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors....
arXiv:2410. 12341v4 Announce Type: replace-cross Abstract: As AI-generated content increasingly populates the web, generative AI models are at growing risk of being trained on their own outputs, a process known as AI autophagy.
arXiv:2606. 05187v1 Announce Type: cross Abstract: Among the many challenges hindering the responsible development and deployment of AI, arguably none has faced more intense scrutiny than bias in its various forms.
arXiv:2602. 16065v2 Announce Type: replace-cross Abstract: As artificial intelligence (AI)-generated content proliferates, models are increasingly trained on their own outputs, risking progressive degradation or collapse.
The paper analyzes the environmental footprint of machine learning model training, focusing on large language models and their hardware. It finds that energy use and environmental impacts have risen exponentially over the past decade, even when employing carbon‑efficient electricity and more efficient hardware. The study argues that optimization strategies alone cannot curb these impacts due to a rebound effect, and stresses the need to evaluate hardware life‑cycle impacts and integrate environmental metrics into NLP research practices.
The paper highlights that as Large Language Models grow in capability and prevalence, their environmental footprint is increasing, yet the machine learning community lacks standardized carbon accounting practices. An automated review of 5,285 NeurIPS 2025 papers shows almost no reporting of environmental impact. To address this, the authors propose standardized sustainability metrics for training efficiency, heuristics for estimating inference carbon costs, a software tool called carbonbenchmark for tracking emissions, and the SMAJ framework to encourage prioritizing computational efficiency and environmental accountability over marginal accuracy gains.
The paper argues that large language models (LLMs) have a growing carbon footprint and that current efficiency gains are offset by rebound effects such as Jevons Paradox. It proposes applying the EU waste hierarchy—prevention, reuse, recycling, recovery, and disposal—to LLMs, suggesting that preventing waste and unnecessary use can significantly reduce environmental impact. The authors emphasize that treating LLMs as products that can become waste offers new ways to motivate more sustainable practices.
arXiv:2511. 06148v4 Announce Type: replace-cross Abstract: As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly important to ensure that they are unbiased.
arXiv:2602. 19789v2 Announce Type: replace Abstract: This position paper argues that the machine learning community must move from preaching to practising data frugality for responsible artificial intelligence (AI) development.
arXiv:2607. 09705v1 Announce Type: cross Abstract: Since 2023, computer scientists have warned against model collapse -- the contamination of training sets with AI-generated outputs that progressively degrade model performance.
arXiv:2603. 11784v2 Announce Type: replace Abstract: As scaling laws push the training of frontier large language models (LLMs) toward ever-growing data requirements, training pipelines are approaching a regime where much of the publicly available online text may be consumed.