arXiv:2606. 02008v1 Announce Type: cross Abstract: Pre-training has become a fundamental paradigm in modern machine learning, with one of its key empirical benefits being reduced downstream sample complexity as the scale of pre-training data increases.
By Kazuto Fukuchi, Ryuichiro Hataya, Kota Matsui
Pre-training has become a fundamental paradigm in modern machine learning, with one of its key empirical benefits being reduced downstream sample complexity as the scale of pre-training data increases. However, existing theoretical frameworks for pre-training do not fully explain this phenomenon.
The paper introduces TESS, a scalable data‑selection framework that replaces per‑sample weights with a selection network to improve transferability across datasets and model sizes. It identifies instability in existing meta‑learning for training‑data selection (MTS) due to weight suppression and overreliance on easy features, and proposes a Pointwise Value Matching objective to address these issues. Experiments on large language model safety and instruction tuning show strong transfer from subsets to full corpora and from smaller to larger models.
By Zilin Du, Bowen Yang, Boyang Albert Li
arXiv:2607. 02850v1 Announce Type: new Abstract: Meta-learning without labeled data is crucial for real-world applications, where obtaining labeled datasets can be expensive or restricted due to privacy concerns.
By Lei Sun, Yusuke Tanaka, Tomoharu Iwata
arXiv:2608. 11746v1 Announce Type: new Abstract: Modern systems are increasingly expected to transfer across tasks not specified during training.
By Ellen Su, Andres Potapczynski, Shikai Qiu, Edward Hughes, Andrew Gordon Wilson
arXiv:2609.09572v1 Announce Type: new
Abstract: Synthetic data has become a promising way to scale model training beyond limited human-generated data but it may also induce strong model collapse (Doh...
By Jichu li, Difan Zou
arXiv:2607. 20465v1 Announce Type: new Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end.
By Hao Liang, Qifeng Cai, Yibo Lin, Jianzhuo Du, Qifeng Xia, Sizhe Qiu, Linzhuang Sun, Meiyi Qiang, Zhaoyang Han, Xiaochen Ma, Bohan Zeng, Ruichuan An, Conghui He, Wentao Zhang
arXiv:2606. 18209v1 Announce Type: new Abstract: Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples.
By Trisha Mittal, Akshay Mehra, Joshua Kimball
The paper compares five machine unlearning (MU) methods—NegGrad, Fine‑Tuning (FT), Random Labeling (RL), SalUn, and MUNBa—on noisy‑label correction across CIFAR‑10, CIFAR‑100, and Food‑101N. Results show that the best MU strategy depends on the noise type: FT works well for most closed‑set noise, RL and SalUn are robust and nearly match retraining accuracy under instance‑dependent noise, while MUNBa excels only under extreme symmetric noise. In open‑set noise, retraining on the cleaned data actually hurts performance, indicating that approximating retraining is not suitable in that regime, yet all MU methods still achieve near‑retraining accuracy on Food‑101N with much lower runtime.
By Jo\~ao L. P. Santana, Filipe R. Cordeiro
The paper introduces FROST, an online framework that filters synthetic training data by estimating its utility through gradient feedback anchored in real data. FROST calibrates batch utility against recent history to decide when to filter, removing 20–30% of synthetic samples while improving performance on image classification and LLM fine-tuning tasks. The method is also applied to a large‑scale industrial ads re‑ranking system, yielding significant gains over an optimized production baseline.
By Yanran Wu, Sana Lakdawala, Renzo Tassara Miller, Chongyang Bai, Sharath Ciddu, Shivendra Pratap Singh, Kungang Li, Sandeep Pandey, Chunwei Liu
arXiv:2608. 09091v1 Announce Type: cross Abstract: Transfer learning is particularly useful in settings with limited training data, and within image classification it is common to transfer learn upon massive datasets like ImageNet , CIFAR-100, or COCO .
By Jing Ning, James D. Braza
arXiv:2506. 01486v2 Announce Type: replace Abstract: Data imbalance persists as a pervasive challenge in regression tasks, introducing bias in model performance and undermining predictive reliability.
By Jelke Wibbeke, Sebastian Rohjans, Andreas Rauh