arXiv Computation and Language

Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research

Researchers use synthetic survey respondents generated by large language models as substitutes for human samples, but current validation methods often compare them to human surveys in ways that may not reflect real-world consequential behaviour. The authors propose a new validation framework that requires explicit statements of how well synthetic data correspond to human behaviour, specifies which diagnostics are addressed, and demands subgroup-level validity claims to avoid misrepresentation. The framework operationalises distributional, procedural, and recognition justice dimensions and introduces within-persona counterfactual experiments, illustrated with a case study on electric vehicle charging tariffs and concluded with a reporting checklist for researchers.

arXiv Computation and Language
Sep 25

Artificial Societies Benchmark: A Validation Framework for Synthetic Research

The article introduces the Artificial Societies Benchmark, a validation framework designed to evaluate synthetic populations used in research. It comprises eleven tests covering internal, construct, and external validity, drawing on twenty human data sources and comparing nine language models. The benchmark links specific research uses to the evidence required and assesses how results vary with different respondent information, revealing that strong performance in one domain does not guarantee fidelity in others.

By Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis, James K. He
arXiv AI
Sep 15

Synthetic Data in Marketing Research: How to Evaluate and When to Trust

The paper discusses the use of synthetic data in marketing research, arguing that the key question is not whether synthetic respondents work, but when they do. It categorizes synthetic data into three types—ungrounded LLM responses, segment-level personas, and individual-level digital twins—and maps each to the decisions they can support. The authors also propose a taxonomy of accuracy measures, highlight the forgotten question problem, and introduce an ex‑ante answerability diagnostic based on R² to improve twin-human correlation.

By Oded Netzer, Rajan Sambandam
arXiv Computation and Language
Sep 23

Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation

The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.

By Alexandre Cristov\~ao Maiorano
arXiv Machine Learning
Sep 3

FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making

FairLens is a benchmark and evaluation framework that measures fairness and validity of vision‑language models (VLMs) in high‑stakes domains such as hiring, legal, and healthcare. It uses over 100,000 face‑image and question pairs covering gender, race, and age, and assesses responses through demographic parity, soundness, demographic association, and bias in free‑text generation. The study finds that VLMs often make unwarranted inferences from faces rather than abstaining, especially in legal and healthcare contexts, and that small parity gaps can still hide unsafe treatment across groups.

By Vahid Reza Khazaie, Ahmed Y. Radwan, Shaina Raza
arXiv AI
Aug 19

Position: Fairness Failure in Generative Models is an Evaluation Problem

The paper argues that fairness failures in generative models arise mainly from inadequate evaluation practices, making fairness findings hard to compare or use for deployment. It diagnoses common empirical and conceptual shortcomings in current methods and calls for a move toward standardized, generative‑specific evaluation. The authors introduce Fairness Cards, a minimal reporting artifact that explicitly documents evaluation choices—such as prompt families, counterfactual protocols, metrics, and refusal handling—to improve reproducibility, comparability, and accountability.

By Mariia Vladimirova, Jean-Yves Franceschi, Thibaut Issenhuth