The paper discusses the use of synthetic data in marketing research, arguing that the key question is not whether synthetic respondents work, but when they do. It categorizes synthetic data into three types—ungrounded LLM responses, segment-level personas, and individual-level digital twins—and maps each to the decisions they can support. The authors also propose a taxonomy of accuracy measures, highlight the forgotten question problem, and introduce an ex‑ante answerability diagnostic based on R² to improve twin-human correlation.
By Oded Netzer, Rajan Sambandam
arXiv:2607. 18164v1 Announce Type: cross Abstract: Digital Twins rely on surrogate models to mirror physical systems in real time, yet these models can degrade as operating conditions evolve, a phenomenon known as concept drift.
By Yi-Ping Chen, Ying-Kuan Tsai, Vispi Karkaria, Seul Lee, Daniel Apley, Wei Chen
arXiv:2608.20344v1 Announce Type: new
Abstract: LLM-based "digital twins" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of...
By Iris Ye, Tianze Deng, Ozan Candogan
arXiv:2606. 04592v1 Announce Type: cross Abstract: LLM-based digital twins promise to scale and accelerate market research, but most published twins are either coarse persona bots conditioned on a few demographic questions or detailed individual-level twins built on purpose-collected surveys and interview transcripts.
By Leonard Kinzinger, Jochen Hartmann
arXiv:2601. 14264v2 Announce Type: replace-cross Abstract: Large language models (LLMs) act as digital twins for human respondents, yet their psychometric comparability remains uncertain.
By Yufei Zhang, Zhihao Ma
arXiv:2301. 07210v5 Announce Type: replace-cross Abstract: Digital twins are simulation-based models designed to predict how a real-world process will evolve in response to interventions.
By Rob Cornish, Muhammad Faaiz Taufiq, Arnaud Doucet, Chris Holmes
arXiv:2606. 27334v1 Announce Type: new Abstract: Digital twins have emerged as a promising paradigm for personalized healthcare, enabling modeling of individual behavior and health trajectories.
By Mohammad Mehdi Hosseini, Mohammad H. Mahoor, Hiroko H. Dodge
Cognitive Digital Twins (CDTs) are dynamic computational models that represent an individual’s cognition, continuously updated with behavioral, contextual, or physiological data to predict or simulate that person’s mental processes or act as a proxy for communication and decision‑making. The paper defines CDTs, distinguishes them from related systems, and introduces a 5A governance framework—authority, autonomy, access & control, accountability, and availability—to address their unique risks. It identifies specific threats such as misrepresentation, epistemic authority shifts, shadow twins, and proxy‑power asymmetries, and proposes governance requirements for high‑risk CDTs, including stronger consent, purpose limitation, validity, traceability, contestation, independent review, and model retirement.
By Vamshi Krishna Bonagiri, Juan Nicolas Sepulveda-Arias, Abdoul Jalil Djiberou Mahamadou, Monojit Choudhury
arXiv:2601. 20819v2 Announce Type: replace-cross Abstract: Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science.
By Yilin Song, Dan M. Kluger, Harsh Parikh, Tian Gu
arXiv:2608.29455v1 Announce Type: cross
Abstract: LLMs are increasingly used as human surrogates, often on the premise that richer persona data could make them substitutes or exploratory tools for sp...
By Daehwan Ahn, Chengfeng Mao, Dokyun Lee
ExploraTwin is an open‑access, non‑profit research platform designed to lower the friction for testing and deploying digital twin simulations. It offers two modes: a survey mode that lets researchers upload or create surveys, select digital twin samples, run simulations, and export analysis‑ready data; and a panel mode that supports small groups of twins for open‑ended conversations, annotation, and moderated voice discussions. The platform also introduced CroissantTwin, a standardized data format for adding twin samples, and demonstrated high fidelity in survey execution with 99.6% valid responses across 197,000 answer units.
By Naveen Venkatanarayanan, Yuchen Qiu, Tianyi Peng, George Gui, Olivier Toubia
arXiv:2607. 16239v1 Announce Type: new Abstract: AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains.
By Lei Shi, Anlan Zhang, Rita Lyu, Zhengmian Hu, Tong Yu, David Arbour, Avi Feller, Saayan Mitra, Ritwik Sinha