arXiv AI

Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure

The paper introduces an adequacy‑aware calibration protocol for generative social simulators that integrates amortized posterior estimation, synthetic identifiability assessment, matched‑sample‑size adequacy checks, diagnosis‑guided repair, and held‑out audits. Applied to a second‑hand luxury resale market, the protocol reveals that behavioural parameters are recoverable but calibration is approximate and overconfident for one parameter, and that the simulator’s reachability reference is violated in every cell, particularly in mean purchased tier. The repair improves two of four cells but fails to restore full adequacy, and a held‑out audit uncovers a buyer‑breadth‑dispersion miss not detected earlier; profile‑source ablation shows language‑model‑derived persona profiles outperform a flat rule baseline, though within‑category brand relabelling has no consistent effect. whyItMatters":"The study demonstrates that without an adequacy check, generative social models may appear valid descriptively yet fail to capture key emergent network structures, highlighting the need for rigorous calibration protocols in social simulation research."

arXiv Computation and Language
Sep 23

Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation

The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.

By Alexandre Cristov\~ao Maiorano
arXiv AI
Jul 24

Benchmarking the Personalization Capabilities of Large Language Models

arXiv:2607. 20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two-party problem in which sender and receiver have independent objectives.

By Ashutosh Srivastava, Siddharth Yedlapati, Vinay Aggarwal, Yaman Kumar Singla, Shashwat Dixit, Jitendra Ajmera, Balaji Krishnamurthy
arXiv AI
Aug 26

AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval

AgentWorld is a simulation framework that evaluates agentic information retrieval by incorporating diverse user personalities based on the Big Five (OCEAN) traits, stateful tool-use environments, and a pass$^k$ consistency metric with structured fault classification and partial-credit scoring. It includes a risk analyzer that uses Monte‑Carlo rollouts and advanced scoring methods to quantify trajectory brittleness and attack attribution. Experiments with conversational analytics, customer‑support agents, and adversarial stress‑testing demonstrate that personality variation reveals failure modes hidden by uniform testing, such as cross‑domain leakage, contextual drift, and significant quality gaps across personas.

By Gunja Agarwal, Arup Kumar Das, Arun Menon, Jitesh Chandra Mishra, Vignesh Divakaran
arXiv Machine Learning
Sep 22

The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation

The paper introduces the Situated Identity Test (SIT), a framework that assesses whether a language model’s behavior can be traced to a specific developmental lineage rather than merely imitating a persona. SIT requires agents to possess accurate knowledge of their recorded experiences while appropriately ignoring ungrounded information, and it demonstrates that policies based only on compressed profiles are limited in distinguishing between colliding life histories. The authors present SITBench, an evaluation suite with 25 profile‑collision pairs and 10,000 probes across nine model architectures, and provide open‑source tools and pilot results on state‑of‑the‑art foundation models.

By Jun He, Deying Yu
arXiv AI
Jun 18

How Well Do Large Language Models Capture Human Personality?

arXiv:2606. 18263v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate human populations via persona prompting, often under the assumptions that richer persona descriptions improve behavioral fidelity, similarly sized attribute combinations are equally simulatable, and persona definitions generalize across tasks.

By Aanisha Bhattacharyya, Yaman Kumar Singla, Rajiv Ratn Shah, Changyou Chen, Jitendra Ajmera
arXiv Machine Learning
Sep 14

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.

By Umesh Bodhwani, Thanh Tran, Kai Wei
arXiv Machine Learning
Aug 5

Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

arXiv:2608. 03842v1 Announce Type: cross Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate.

By Nathan Labiosa, David Buff, Ena Nayak, Erica Donno