arXiv AI By Tengfei Shao, Chao Li, Xu Wang, Masayuki Goto

Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure

Read the original on arXiv AI →

The paper introduces an adequacy‑aware calibration protocol for generative social simulators that integrates amortized posterior estimation, synthetic identifiability assessment, matched‑sample‑size adequacy checks, diagnosis‑guided repair, and held‑out audits. Applied to a second‑hand luxury resale market, the protocol reveals that behavioural parameters are recoverable but calibration is approximate and overconfident for one parameter, and that the simulator’s reachability reference is violated in every cell, particularly in mean purchased tier. The repair improves two of four cells but fails to restore full adequacy, and a held‑out audit uncovers a buyer‑breadth‑dispersion miss not detected earlier; profile‑source ablation shows language‑model‑derived persona profiles outperform a flat rule baseline, though within‑category brand relabelling has no consistent effect. whyItMatters":"The study demonstrates that without an adequacy check, generative social models may appear valid descriptively yet fail to capture key emergent network structures, highlighting the need for rigorous calibration protocols in social simulation research."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 23

Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation

The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.

By Alexandre Cristov\~ao Maiorano
arXiv AI
Jul 24

Benchmarking the Personalization Capabilities of Large Language Models

arXiv:2607. 20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two-party problem in which sender and receiver have independent objectives.

By Ashutosh Srivastava, Siddharth Yedlapati, Vinay Aggarwal, Yaman Kumar Singla, Shashwat Dixit, Jitendra Ajmera, Balaji Krishnamurthy
arXiv AI
Aug 26

AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval

AgentWorld is a simulation framework that evaluates agentic information retrieval by incorporating diverse user personalities based on the Big Five (OCEAN) traits, stateful tool-use environments, and a pass$^k$ consistency metric with structured fault classification and partial-credit scoring. It includes a risk analyzer that uses Monte‑Carlo rollouts and advanced scoring methods to quantify trajectory brittleness and attack attribution. Experiments with conversational analytics, customer‑support agents, and adversarial stress‑testing demonstrate that personality variation reveals failure modes hidden by uniform testing, such as cross‑domain leakage, contextual drift, and significant quality gaps across personas.

By Gunja Agarwal, Arup Kumar Das, Arun Menon, Jitesh Chandra Mishra, Vignesh Divakaran