arXiv:2509. 20345v3 Announce Type: replace-cross Abstract: The rapid proliferation of high-quality synthetic data -- generated by advanced AI models or collected as auxiliary data from related tasks -- presents both opportunities and challenges for statistical inference.
By Meshi Bashari, Yonghoon Lee, Roy Maor Lotan, Edgar Dobriban, Yaniv Romano
The paper introduces a framework for synthetic‑augmented inference that balances the number of synthetic observations with their assigned weight. It defines a size‑weight frontier, estimating for each weight the maximum synthetic sample size that still guarantees target task‑marginal coverage for all smaller sizes. The authors provide finite‑sample coverage guarantees for configurations on or below this frontier and demonstrate that, when applied to augment opinion survey data with large language model responses, the method achieves the desired coverage while significantly tightening confidence intervals.
By Chengpiao Huang, Kaizheng Wang
arXiv:2601. 17717v3 Announce Type: replace Abstract: Large Language Models (LLMs) have emerged as powerful tools for generating data across various modalities.
By Kaituo Zhang, Mingzhi Hu, Hoang Anh Duy Le, Fariha Kabir Torsha, Zhimeng Jiang, Minh Khai Bui, Chia-Yuan Chang, Yu-Neng Chuang, Zhen Xiong, Ying Lin, Guanchu Wang, Na Zou
arXiv:2604. 17267v2 Announce Type: replace Abstract: Large Language Models can generate synthetic survey responses at low cost, but their accuracy varies unpredictably across questions.
By Zikun Ye, Hema Yoganarasimhan
arXiv:2603. 00059v3 Announce Type: replace-cross Abstract: How well can AI-derived synthetic research data replicate the responses of human participants?
By Jason Miklian, Kristian Hoelscher, John E. Katsos
The paper investigates how synthetic data generated by large language models (LLMs) can be characterized using sample-level learnability derived from encoder training dynamics. It compares different LLM families and scales across tasks such as single- and multi-label classification, labeling, and tree prediction, and contrasts these synthetic datasets with human-written data. The study also examines the robustness of learned data distributions across encoders and evaluates how data selection strategies based on learnability signals impact the performance of both synthetic and organic data.
By Irene Lago, Ana Ezquerro, David Vilares