arXiv AI

Valid Inference with Synthetic Data via Task Exchangeability

arXiv:2606. 13629v1 Announce Type: cross Abstract: There is a proliferation of work arguing for the use of synthetic data in scientific research.

arXiv Machine Learning
Jun 5

General Synthetic-Powered Inference

arXiv:2509. 20345v3 Announce Type: replace-cross Abstract: The rapid proliferation of high-quality synthetic data -- generated by advanced AI models or collected as auxiliary data from related tasks -- presents both opportunities and challenges for statistical inference.

By Meshi Bashari, Yonghoon Lee, Roy Maor Lotan, Edgar Dobriban, Yaniv Romano
arXiv Machine Learning
Aug 31

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

The paper introduces a framework for synthetic‑augmented inference that balances the number of synthetic observations with their assigned weight. It defines a size‑weight frontier, estimating for each weight the maximum synthetic sample size that still guarantees target task‑marginal coverage for all smaller sizes. The authors provide finite‑sample coverage guarantees for configurations on or below this frontier and demonstrate that, when applied to augment opinion survey data with large language model responses, the method achieves the desired coverage while significantly tightening confidence intervals.

By Chengpiao Huang, Kaizheng Wang
arXiv Computation and Language
3d ago

Synthetic Data Characterization via Training Dynamics

The paper investigates how synthetic data generated by large language models (LLMs) can be characterized using sample-level learnability derived from encoder training dynamics. It compares different LLM families and scales across tasks such as single- and multi-label classification, labeling, and tree prediction, and contrasts these synthetic datasets with human-written data. The study also examines the robustness of learned data distributions across encoders and evaluates how data selection strategies based on learnability signals impact the performance of both synthetic and organic data.

By Irene Lago, Ana Ezquerro, David Vilares
arXiv AI
Jun 11

Can AI Agents Synthesize Scientific Conclusions?

arXiv:2606. 11337v1 Announce Type: new Abstract: Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions.

By Hayoung Jung, Pedro Viana Diniz, Jos\'e Reinaldo Corr\^ea Roveda, Abner Fernandes da Silva, Haeun Jung, Enoch Tsai, Aleksandra Korolova, Manoel Horta Ribeiro
arXiv Machine Learning
2d ago

Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data

GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.

By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar