Valid Inference with Synthetic Data via Task Exchangeability
arXiv:2606. 13629v1 Announce Type: cross Abstract: There is a proliferation of work arguing for the use of synthetic data in scientific research.
arXiv:2509. 20345v3 Announce Type: replace-cross Abstract: The rapid proliferation of high-quality synthetic data -- generated by advanced AI models or collected as auxiliary data from related tasks -- presents both opportunities and challenges for statistical inference.
arXiv:2606. 13629v1 Announce Type: cross Abstract: There is a proliferation of work arguing for the use of synthetic data in scientific research.
The paper introduces a framework for synthetic‑augmented inference that balances the number of synthetic observations with their assigned weight. It defines a size‑weight frontier, estimating for each weight the maximum synthetic sample size that still guarantees target task‑marginal coverage for all smaller sizes. The authors provide finite‑sample coverage guarantees for configurations on or below this frontier and demonstrate that, when applied to augment opinion survey data with large language model responses, the method achieves the desired coverage while significantly tightening confidence intervals.
The paper introduces a conformal prediction framework designed for molecular property prediction under label shift. By weighting conformal scores with marginal label probability ratios, it generates statistically rigorous prediction intervals without retraining, enabling robust uncertainty quantification when property distributions change. This approach provides actionable confidence measures that improve the reliability of AI-driven predictions in drug discovery.
arXiv:2610.00814v1 Announce Type: cross Abstract: Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data m...
arXiv:2606. 06724v1 Announce Type: new Abstract: Representative data is fundamental in machine learning, as limited data hinders generalisation.
GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.
arXiv:2606. 00563v1 Announce Type: cross Abstract: Selection bias is a common and often unavoidable aspect of real-world data that challenges the generalizability of machine learning models.
arXiv:2607. 08347v1 Announce Type: cross Abstract: Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled.
arXiv:2601. 20819v2 Announce Type: replace-cross Abstract: Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science.
arXiv:2609.24629v1 Announce Type: new Abstract: A/B testing requires large sample sizes, long timelines, and significant costs. When auxiliary predictions of experimental outcomes are available from...
arXiv:2602. 24007v3 Announce Type: replace-cross Abstract: Protein function relies on dynamic conformational ensembles, yet current generative models like AlphaFold3 often fail to produce ensembles that match experimental data.
arXiv:2607. 09404v1 Announce Type: new Abstract: Motivation: Rare disease (RD) diagnosis is frequently delayed due to the similarities in symptoms to common disease variants.