Hugging Face Blog

Introducing the Synthetic Data Generator - Build Datasets with Natural Language

arXiv Machine Learning
Jul 31

How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data

arXiv:2604. 13977v2 Announce Type: replace-cross Abstract: Synthetic data is a standard component in training large language models, yet systematic comparisons across design dimensions, including rephrasing strategy, generator model, and source data, remain absent.

By Joel Niklaus, Atsuki Yamaguchi, Michal \v{S}tef\'anik, Guilherme Penedo, Hynek Kydl\'i\v{c}ek, Elie Bakouch, Lewis Tunstall, Edward Emanuel Beeching, Thibaud Frere, Colin Raffel, Leandro von Werra, Thomas Wolf
arXiv Machine Learning
Jul 7

Memisis: Orchestrating and Evaluating Synthetic Data for Tabular Health Datasets

arXiv:2605. 17758v2 Announce Type: replace Abstract: Synthetic data is widely used in healthcare to create datasets that preserve statistical properties of real data without exposing sensitive patient information.

By Nitish Nagesh, Pengbao Zhou, Atchuth Naveen Chilaparasetti, Yajat Nagaraj Kiran, Tu Nguyen, Arshia Harish Puthran, Muhjaazee Love, Aadi Sharma, Mahdi Bagheri, Ian Harris, Amir M. Rahmani
arXiv Machine Learning
Jun 9

Disjoint Generation of Synthetic Data

arXiv:2507. 19700v2 Announce Type: replace Abstract: We propose a new framework for generating tabular synthetic datasets via disjoint generative models.

By Anton Danholt Lautrup, Muhammad Rajabinasab, Tobias Hyrup, Arthur Zimek, Peter Schneider-Kamp
arXiv Machine Learning
Jul 15

Hierarchical Synthetic Tabular Data Generation: A Hybrid Top-Down and Bottom-Up Framework

arXiv:2605. 28198v2 Announce Type: replace Abstract: Existing approaches for synthetic tabular data generation are based on either purely generative models or LLMs, both of which struggle with data heterogeneity, logical consistency, rare-event coverage, and robustness in low-data regimes.

By Junfeng Nie, Alvin Jin, Xiaohui Chen