arXiv Computation and Language By Irene Lago, Ana Ezquerro, David Vilares

Synthetic Data Characterization via Training Dynamics

Read the original on arXiv Computation and Language →

The paper investigates how synthetic data generated by large language models (LLMs) can be characterized using sample-level learnability derived from encoder training dynamics. It compares different LLM families and scales across tasks such as single- and multi-label classification, labeling, and tree prediction, and contrasts these synthetic datasets with human-written data. The study also examines the robustness of learned data distributions across encoders and evaluates how data selection strategies based on learnability signals impact the performance of both synthetic and organic data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Sep 25

Let Training Guide Selection: Online Synthetic Data Filtering via Real-Anchored Utility

The paper introduces FROST, an online framework that filters synthetic training data by estimating its utility through gradient feedback anchored in real data. FROST calibrates batch utility against recent history to decide when to filter, removing 20–30% of synthetic samples while improving performance on image classification and LLM fine-tuning tasks. The method is also applied to a large‑scale industrial ads re‑ranking system, yielding significant gains over an optimized production baseline.

By Yanran Wu, Sana Lakdawala, Renzo Tassara Miller, Chongyang Bai, Sharath Ciddu, Shivendra Pratap Singh, Kungang Li, Sandeep Pandey, Chunwei Liu
arXiv Machine Learning
Jul 30

The Advantage of Fine-Grained Training

arXiv:2509. 05130v2 Announce Type: replace Abstract: In classification problems, models are trained to predict a class label based on the input data features.

By Davide Pirovano, Federico Milanesio, Michele Caselle, Piero Fariselli, Matteo Osella
arXiv AI
Sep 10

Generating Pretraining Tokens from Organic Data for Data-Bound Scaling

The paper introduces SynPro, a synthetic data generation framework that augments limited organic text for large language model pretraining by applying rephrasing and reformatting operations. SynPro’s generators are optimized with reinforcement learning rewards for quality, faithfulness, and data influence, and are updated continuously as training plateaus. Experiments on 400M, 1.1B, and 2B models show that SynPro can unlock 3.4–5.2× the effective tokens of standard repetition, even outperforming a non‑data‑bound oracle at larger scales.

By Zichun Yu, Chenyan Xiong