arXiv:2606. 04592v1 Announce Type: cross Abstract: LLM-based digital twins promise to scale and accelerate market research, but most published twins are either coarse persona bots conditioned on a few demographic questions or detailed individual-level twins built on purpose-collected surveys and interview transcripts.
By Leonard Kinzinger, Jochen Hartmann
The paper discusses the use of synthetic data in marketing research, arguing that the key question is not whether synthetic respondents work, but when they do. It categorizes synthetic data into three types—ungrounded LLM responses, segment-level personas, and individual-level digital twins—and maps each to the decisions they can support. The authors also propose a taxonomy of accuracy measures, highlight the forgotten question problem, and introduce an ex‑ante answerability diagnostic based on R² to improve twin-human correlation.
By Oded Netzer, Rajan Sambandam
arXiv:2609.07987v1 Announce Type: new
Abstract: LLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide lit...
By Steven Wang, Kyle Hunt, Shaojie Tang, Kenneth Joseph
arXiv:2603. 00059v3 Announce Type: replace-cross Abstract: How well can AI-derived synthetic research data replicate the responses of human participants?
By Jason Miklian, Kristian Hoelscher, John E. Katsos
arXiv:2601. 14264v2 Announce Type: replace-cross Abstract: Large language models (LLMs) act as digital twins for human respondents, yet their psychometric comparability remains uncertain.
By Yufei Zhang, Zhihao Ma
arXiv:2608.20344v1 Announce Type: new
Abstract: LLM-based "digital twins" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of...
By Iris Ye, Tianze Deng, Ozan Candogan
arXiv:2607. 26060v1 Announce Type: cross Abstract: LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment.
By Cristovao Iglesias, Devesh Batra, Alankar Atreya, Stefan Wagner, Robert Hankache, Patrick Sinclair, Giulio Pelosio, Michael McMillan, Greig A. Cowan, Raad Khraishi
ExploraTwin is an open‑access, non‑profit research platform designed to lower the friction for testing and deploying digital twin simulations. It offers two modes: a survey mode that lets researchers upload or create surveys, select digital twin samples, run simulations, and export analysis‑ready data; and a panel mode that supports small groups of twins for open‑ended conversations, annotation, and moderated voice discussions. The platform also introduced CroissantTwin, a standardized data format for adding twin samples, and demonstrated high fidelity in survey execution with 99.6% valid responses across 197,000 answer units.
By Naveen Venkatanarayanan, Yuchen Qiu, Tianyi Peng, George Gui, Olivier Toubia
arXiv:2509. 02910v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly act on people's behalf: they write emails, buy groceries, and book restaurants.
By Sandra C. Matz, Kimberly Klugescheid, C. Blaine Horton, Sofie Goethals
arXiv:2607. 11039v1 Announce Type: cross Abstract: Young job seekers frequently turn to social media to compare themselves with peers and make sense of career possibilities.
By Pengping Tan, Baoquan Zhao, Zhenhui Peng
arXiv:2606. 16475v1 Announce Type: cross Abstract: Many societal decisions are settled by contests of persuasion.
By Kobi Hackenburg, Caroline Wagner, Luke Hewitt, Ben M. Tappin, Ed Saunders, Hannah Rose Kirk, Helen Margetts, Christopher Summerfield
The paper investigates whether human psychometric questionnaires can reliably characterize large language models (LLMs) in everyday interactions. By comparing eight open‑source LLMs’ value and personality profiles from Likert self‑reports (PVQ‑40/21 and BFI‑44/10) with generation probabilities on value‑laden user queries, the authors find substantial divergence between the two methods. The study shows that questionnaire items contain explicit lexical cues that lead models to respond in socially desirable ways, whereas realistic user queries lack such cues, and demographic persona prompts shift questionnaire responses but not generation outputs, indicating that questionnaire scores overestimate LLMs’ true behavioral tendencies.
By Woojung Song, Dongmin Choi, Yoonah Park, Jongwook Han, Eun-Ju Lee, Yohan Jo