The paper discusses the use of synthetic data in marketing research, arguing that the key question is not whether synthetic respondents work, but when they do. It categorizes synthetic data into three types—ungrounded LLM responses, segment-level personas, and individual-level digital twins—and maps each to the decisions they can support. The authors also propose a taxonomy of accuracy measures, highlight the forgotten question problem, and introduce an ex‑ante answerability diagnostic based on R² to improve twin-human correlation.
By Oded Netzer, Rajan Sambandam
AI‑moderated interviews are a scalable market‑research method that can match human moderation in depth, cover more themes, and recover more customer needs while keeping budgets constant. Participants, however, feel more emotionally engaged with live humans. Digital twins built from AI‑moderated data predict consumer responses better than demographics‑only personas, but the added richness does not improve quantitative predictions over static interviews, and prediction errors stem from differences in thinking styles and data‑distribution gaps.
By Yuting Deng, Jingxuan Liu, Olivier Toubia, Naman Jain
arXiv:2608.20344v1 Announce Type: new
Abstract: LLM-based "digital twins" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of...
By Iris Ye, Tianze Deng, Ozan Candogan
arXiv:2609.07987v1 Announce Type: new
Abstract: LLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide lit...
By Steven Wang, Kyle Hunt, Shaojie Tang, Kenneth Joseph
arXiv:2601. 14264v2 Announce Type: replace-cross Abstract: Large language models (LLMs) act as digital twins for human respondents, yet their psychometric comparability remains uncertain.
By Yufei Zhang, Zhihao Ma
arXiv:2609.13148v1 Announce Type: cross
Abstract: Large language models are increasingly deployed as synthetic consumer panels, promising $97\%$ cost reductions over traditional surveys. Yet aggregat...
By Robson Tigre, Hugo Gobato Souto
ExploraTwin is an open‑access, non‑profit research platform designed to lower the friction for testing and deploying digital twin simulations. It offers two modes: a survey mode that lets researchers upload or create surveys, select digital twin samples, run simulations, and export analysis‑ready data; and a panel mode that supports small groups of twins for open‑ended conversations, annotation, and moderated voice discussions. The platform also introduced CroissantTwin, a standardized data format for adding twin samples, and demonstrated high fidelity in survey execution with 99.6% valid responses across 197,000 answer units.
By Naveen Venkatanarayanan, Yuchen Qiu, Tianyi Peng, George Gui, Olivier Toubia
arXiv:2607. 26348v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions.
By Zihan Chen, Di Zhu, Lei Nico Zheng
The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.
By Alexandre Cristov\~ao Maiorano
arXiv:2607. 18310v1 Announce Type: cross Abstract: Synthetic-population tools increasingly run every individual as an independent large language model (LLM) agent.
By Gurkan Ozkan
arXiv:2608. 14606v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level.
By Mantas Lukauskas, Viktorija \v{S}arkauskait\.e
arXiv:2608.30023v1 Announce Type: cross
Abstract: Generative engines such as ChatGPT, Gemini, and Perplexity answer buyer questions directly and name a shortlist of brands inside the answer. Studying...
By Dmitrij \.Zatuchin, Daniil Dzemesjuk