arXiv:2609.07987v1 Announce Type: new
Abstract: LLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide lit...
By Steven Wang, Kyle Hunt, Shaojie Tang, Kenneth Joseph
arXiv:2606. 00563v1 Announce Type: cross Abstract: Selection bias is a common and often unavoidable aspect of real-world data that challenges the generalizability of machine learning models.
By Kara Liu, Maggie Wang, Russ B. Altman
arXiv:2601. 20819v2 Announce Type: replace-cross Abstract: Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science.
By Yilin Song, Dan M. Kluger, Harsh Parikh, Tian Gu
arXiv:2609.35815v1 Announce Type: cross
Abstract: Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated co...
By Ian Arawjo
arXiv:2604. 14575v2 Announce Type: replace-cross Abstract: Large language models enable inexpensive AI-generated annotations, but using them reliably for causal inference remains challenging.
By Cheng Lu, Mengxin Wang, Dennis J. Zhang, Heng Zhang
arXiv:2608. 16196v1 Announce Type: new Abstract: Personalized game generation requires inferring a player's abilities and behavioral style from how they play.
By Yifan Lu, Xiaopeng Yuan, Haohan Wang
arXiv:2608. 07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims.
By Jiacheng Miao, Jin Mu, Guanhua Chen, James Zou
arXiv:2509.24988v2 Announce Type: replace-cross
Abstract: Generating accurate and calibrated confidence estimates is critical for deploying LLMs in high-stakes or user-facing applications, and remain...
By Hanqi Xiao, Vaidehi Patil, Hyunji Lee, Elias Stengel-Eskin, Mohit Bansal
The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.
By Alexandre Cristov\~ao Maiorano
arXiv:2411. 10109v3 Announce Type: replace Abstract: Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes.
By Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein
arXiv:2507. 02169v2 Announce Type: replace Abstract: Machine learning models are often used in applications where their inputs change due to routine interactions, strategic manipulation, or noise.
By Harry Cheon, Meredith Stewart, Bogdan Kulynych, Tsui-Wei Weng, Berk Ustun
arXiv:2606. 29784v1 Announce Type: cross Abstract: Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity.
By Xinrui Ruan, Zhenyu Zhao, Waverly Wei, Yueshan Zhang, Zeyu Zheng, Sui Huang, Jingshen Wang