Epistemic diversity across language models mitigates knowledge collapse
arXiv:2512. 15011v3 Announce Type: replace Abstract: Artificial intelligence (AI) increasingly generates the very content used to train future AI systems.
arXiv:2410. 12341v4 Announce Type: replace-cross Abstract: As AI-generated content increasingly populates the web, generative AI models are at growing risk of being trained on their own outputs, a process known as AI autophagy.
arXiv:2512. 15011v3 Announce Type: replace Abstract: Artificial intelligence (AI) increasingly generates the very content used to train future AI systems.
arXiv:2510. 01171v4 Announce Type: replace-cross Abstract: Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collapse.
arXiv:2509. 09960v2 Announce Type: replace-cross Abstract: Synthetic tabular data generation is increasingly essential in machine learning, supporting downstream applications when real-world, high-quality tabular data is insufficient.
arXiv:2608.21366v1 Announce Type: new Abstract: Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors....
arXiv:2505. 20161v2 Announce Type: replace-cross Abstract: Effective generalization in language models depends critically on the diversity of their training data.
arXiv:2604. 24927v2 Announce Type: replace-cross Abstract: Generating diverse responses is crucial for test-time scaling of large language models (LLMs), yet standard stochastic sampling mostly yields surface-level lexical variation, limiting semantic exploration.
arXiv:2603. 11784v2 Announce Type: replace Abstract: As scaling laws push the training of frontier large language models (LLMs) toward ever-growing data requirements, training pipelines are approaching a regime where much of the publicly available online text may be consumed.
arXiv:2609.15094v1 Announce Type: cross Abstract: In industrial recommendation feeds, presenting a static headline for an item often fails to satisfy the diverse, multimodal interests of the user pop...
arXiv:2609.37891v1 Announce Type: cross Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--...
arXiv:2606. 04928v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed across diverse applications, raising critical questions for governance, accountability, and data provenance.
arXiv:2605. 10574v3 Announce Type: replace Abstract: As artificial intelligence advances, models are not improving uniformly.
arXiv:2602. 03300v2 Announce Type: replace-cross Abstract: In this work, we aim to develop effective data synthesis techniques that autonomously synthesize multimodal training data for enhancing MLLMs in solving complex real-world tasks.