arXiv:2610.00814v1 Announce Type: cross
Abstract: Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data m...
By Yang Ba, Michelle V. Mancenido, Rong Pan
arXiv:2605. 09697v3 Announce Type: replace-cross Abstract: In many real-world computer vision applications, including medical imaging and industrial inspection, binary classification tasks are characterized by a severe scarcity of positive samples.
By Radhika Amar Desai, Modigari Narendra
arXiv:2607. 02637v1 Announce Type: cross Abstract: Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models.
By Disheng Liu, Tuo Liang, Chaoda Song, Yu Yin
The paper introduces the Real‑Calibrated Synthetic‑First Data Engine, a modular pipeline that integrates controllable diffusion‑based synthetic image generation with multi‑stage curation, filtering, and optional uncertainty‑driven selection and human verification. Designed as a CLI‑based framework, it allows independent configuration of generation, filtering, selection, and validation modules to enhance reproducibility and flexibility in real‑world data workflows. Empirical tests on human pose estimation demonstrate that synthetic data can boost a real‑data baseline when used as low‑cost augmentation, though synthetic‑only training still lags behind real‑only performance, underscoring the importance of data‑centric orchestration in low‑data regimes.
By Yukang Shen, Zhiguo Liu, Yingshu Li, Yan Huang
arXiv:2308. 04553v4 Announce Type: replace-cross Abstract: Visual recognition models are prone to learning spurious correlations induced by a biased training set where certain conditions $B$ (\eg, Indoors) are over-represented in certain classes $Y$ (\eg, Big Dogs).
By Maan Qraitem, Kate Saenko, Bryan A. Plummer
arXiv:2602.05391v3 Announce Type: replace
Abstract: Dataset distillation seeks to synthesize a compact surrogate dataset that enables performance comparable to training on the original dataset for do...
By Qianxin Xia, Jiawei Du, Yuhan Zhang, Xin Zhang, Xuewan He, Wenbo Jiang, Jielei Wang, Tao Luo, Guoming Lu
arXiv:2609.38476v1 Announce Type: new
Abstract: Synthetic data are most valuable where general-purpose datasets cannot provide the domain-specific priors a task requires, and where manual annotation...
By Saptarshi Neil Sinha, Paul Julius K\"uhn, Michael Weinmann
arXiv:2606. 00571v1 Announce Type: cross Abstract: Synthetic data are increasingly used to train neural networks, yet distributional mismatch with real data limits their effectiveness when used indiscriminately.
By Zilin Du, Junqi Zhao, Boyang Albert Li
arXiv:2606. 06458v1 Announce Type: new Abstract: Multiple Instance Learning (MIL) addresses problems where supervision is available at the level of bags of instances and has been successfully applied in fields ranging from computational pathology to satellite imagery.
By Alexander M\"ollers, Marvin Sextro, Julius Hense, Gabriel Dernbach, Klaus-Robert M\"uller
arXiv:2509. 05130v2 Announce Type: replace Abstract: In classification problems, models are trained to predict a class label based on the input data features.
By Davide Pirovano, Federico Milanesio, Michele Caselle, Piero Fariselli, Matteo Osella
arXiv:2502. 18049v5 Announce Type: replace-cross Abstract: Recent studies identified an intriguing phenomenon in recursive generative model training known as model collapse, where models trained on data generated by previous models exhibit severe performance degradation.
By Hengzhi He, Shirong Xu, Guang Cheng
arXiv:2402.11215v4 Announce Type: replace
Abstract: The choice of batch size in minibatch stochastic gradient optimization is critical for both optimization and generalization performance in large-sc...
By Tim Tsz-Kit Lau, Han Liu, Mladen Kolar