arXiv AI

Synthetic Data in Marketing Research: How to Evaluate and When to Trust

The paper discusses the use of synthetic data in marketing research, arguing that the key question is not whether synthetic respondents work, but when they do. It categorizes synthetic data into three types—ungrounded LLM responses, segment-level personas, and individual-level digital twins—and maps each to the decisions they can support. The authors also propose a taxonomy of accuracy measures, highlight the forgotten question problem, and introduce an ex‑ante answerability diagnostic based on R² to improve twin-human correlation.

arXiv AI
Jun 4

Synthetic Personalities: How Well Can LLMs Mimic Individual Respondents Using Socio-Economic Microdata?

arXiv:2606. 04592v1 Announce Type: cross Abstract: LLM-based digital twins promise to scale and accelerate market research, but most published twins are either coarse persona bots conditioned on a few demographic questions or detailed individual-level twins built on purpose-collected surveys and interview transcripts.

By Leonard Kinzinger, Jochen Hartmann
arXiv AI
Jul 7

Silicon Sampling via Cross-Survey Transfer

arXiv:2607. 03091v1 Announce Type: new Abstract: Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research.

By Chan-Tung Ku, Chan Hsu, Pei-Cing Huang, Frank Cheng-shan Liu, I-Ling Cheng, Yihuang Kang