arXiv Computation and Language By Florian Kutzner, Celina Kacperski, Laura de Moli\`ere, Edoardo Chidichimo, Min Jun Jung, Felix Patrick Sedgwick Wallis, James Kunling He

Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research

Read the original on arXiv Computation and Language →

Researchers use synthetic survey respondents generated by large language models as substitutes for human samples, but current validation methods often compare them to human surveys in ways that may not reflect real-world consequential behaviour. The authors propose a new validation framework that requires explicit statements of how well synthetic data correspond to human behaviour, specifies which diagnostics are addressed, and demands subgroup-level validity claims to avoid misrepresentation. The framework operationalises distributional, procedural, and recognition justice dimensions and introduces within-persona counterfactual experiments, illustrated with a case study on electric vehicle charging tariffs and concluded with a reporting checklist for researchers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 25

Artificial Societies Benchmark: A Validation Framework for Synthetic Research

The article introduces the Artificial Societies Benchmark, a validation framework designed to evaluate synthetic populations used in research. It comprises eleven tests covering internal, construct, and external validity, drawing on twenty human data sources and comparing nine language models. The benchmark links specific research uses to the evidence required and assesses how results vary with different respondent information, revealing that strong performance in one domain does not guarantee fidelity in others.

By Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis, James K. He
arXiv AI
Sep 15

Synthetic Data in Marketing Research: How to Evaluate and When to Trust

The paper discusses the use of synthetic data in marketing research, arguing that the key question is not whether synthetic respondents work, but when they do. It categorizes synthetic data into three types—ungrounded LLM responses, segment-level personas, and individual-level digital twins—and maps each to the decisions they can support. The authors also propose a taxonomy of accuracy measures, highlight the forgotten question problem, and introduce an ex‑ante answerability diagnostic based on R² to improve twin-human correlation.

By Oded Netzer, Rajan Sambandam