arXiv Computation and Language By Howard Kim, Keun Tae Cho

Distributional Validity and Calibration of a Korean Synthetic Persona Panel for Digital and AI Service Use: A Secondary-Data Validation Against the Korea Media Panel Survey

Read the original on arXiv Computation and Language →

The study evaluates a Korean synthetic persona panel (NVIDIA Nemotron‑Personas‑Korea) conditioned on Gemini 3.5 Flash and EXAONE against the KISDI Korea Media Panel Survey. Across eight digital‑AI service‑use indicators and eight innovativeness constructs, the panels achieved mean absolute errors of 15–19 pp, with segment‑level errors up to 52.4 pp and correlation coefficients between 0.69 and 0.90. Holdout calibration reduced sex‑by‑age cell errors but still lagged behind direct real‑data estimation, indicating that synthetic panels are useful diagnostically but not as survey substitutes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 10

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

The paper introduces a three-tier persona vector for user simulation in evaluating LLM agents, comprising 23 dimensions across demographics, behavioral traits, and emotional states, plus a query-complexity overlay. It demonstrates that these nuanced personas generate diverse, scenario-reactive conversations, leading to significant variations in agent goal achievement and compliance across different contexts. The model’s design allows for reproducible, auditable user behavior patterns without relying on learned covariance matrices.

By Rahul Khedar, Eshita, Sneha Teja Sree Reddy Thondapu, Mayank Malhotra, Arup Kumar Das, Jitesh Chandra Mishra, Arun Menon, Avinash Karn, Mouli V
Hugging Face Trending Papers
Sep 8

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

The paper introduces a three-tier persona vector to generate diverse, realistic user inputs for evaluating tool-augmented LLM agents. The vector includes 23 dimensions: categorical demographics, continuous behavioral traits, and continuous emotional states, plus a query-complexity overlay. Experiments on 64,698 conversations show that these persona dimensions produce measurable differences in agent performance and realistic scenario-reactive behavior.