arXiv AI By Lezhi Tan, Tijana Zrnic

Valid Inference with Synthetic Data via Task Exchangeability

Read the original on arXiv AI →

arXiv:2606. 13629v1 Announce Type: cross Abstract: There is a proliferation of work arguing for the use of synthetic data in scientific research.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 5

General Synthetic-Powered Inference

arXiv:2509. 20345v3 Announce Type: replace-cross Abstract: The rapid proliferation of high-quality synthetic data -- generated by advanced AI models or collected as auxiliary data from related tasks -- presents both opportunities and challenges for statistical inference.

By Meshi Bashari, Yonghoon Lee, Roy Maor Lotan, Edgar Dobriban, Yaniv Romano
arXiv Machine Learning
Aug 31

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

The paper introduces a framework for synthetic‑augmented inference that balances the number of synthetic observations with their assigned weight. It defines a size‑weight frontier, estimating for each weight the maximum synthetic sample size that still guarantees target task‑marginal coverage for all smaller sizes. The authors provide finite‑sample coverage guarantees for configurations on or below this frontier and demonstrate that, when applied to augment opinion survey data with large language model responses, the method achieves the desired coverage while significantly tightening confidence intervals.

By Chengpiao Huang, Kaizheng Wang
arXiv Computation and Language
3d ago

Synthetic Data Characterization via Training Dynamics

The paper investigates how synthetic data generated by large language models (LLMs) can be characterized using sample-level learnability derived from encoder training dynamics. It compares different LLM families and scales across tasks such as single- and multi-label classification, labeling, and tree prediction, and contrasts these synthetic datasets with human-written data. The study also examines the robustness of learned data distributions across encoders and evaluates how data selection strategies based on learnability signals impact the performance of both synthetic and organic data.

By Irene Lago, Ana Ezquerro, David Vilares