arXiv AI By Zikun Ye, Hema Yoganarasimhan

Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys

Read the original on arXiv AI →

arXiv:2604. 17267v2 Announce Type: replace Abstract: Large Language Models can generate synthetic survey responses at low cost, but their accuracy varies unpredictably across questions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 31

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

The paper introduces a framework for synthetic‑augmented inference that balances the number of synthetic observations with their assigned weight. It defines a size‑weight frontier, estimating for each weight the maximum synthetic sample size that still guarantees target task‑marginal coverage for all smaller sizes. The authors provide finite‑sample coverage guarantees for configurations on or below this frontier and demonstrate that, when applied to augment opinion survey data with large language model responses, the method achieves the desired coverage while significantly tightening confidence intervals.

By Chengpiao Huang, Kaizheng Wang
arXiv AI
Sep 24

Fine-Tune, Then Rectify

The paper proposes a two‑stage framework that first fine‑tunes a large language model (LLM) and then rectifies its outputs, allocating limited labeled data optimally between the stages. It argues that the usual mean‑squared‑error objective for fine‑tuning misaligns with the downstream rectification, and instead suggests minimizing prediction‑error variance for mean estimation or a scalarized variance metric for general M‑estimation. Empirical results confirm that this variance‑based fine‑tuning, combined with optimal data allocation, yields significant efficiency gains over using either fine‑tuning or rectification alone, or using the conventional objective.

By Zikun Ye, Jinglong Zhao, Lei Wang