Wiki-Talkie: Multilingual Benchmarking of Persona-Based Agents on Real-World Discussions
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.22607v1 Announce Type: new Abstract: We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas,...
arXiv:2608.30873v1 Announce Type: cross Abstract: LLMs are increasingly used for interpersonal advice and as tools for studying social behavior across languages and cultures. A common shortcut for el...
The study introduces a World Values Survey–grounded simulation framework to test whether large language model agents can faithfully represent diverse human value systems. In about 4,000 conversations with 1,200 personas across three models, more than half of the agents failed to express their assigned value profiles from the start, and only 2–7% drifted over time. The results show systematic deviations from the intended value distributions and reveal that simulated dialogues differ from human discussions in their balance of stylistic consistency and semantic diversity.
The paper introduces a Situation–Internal state–Behavior Persona method to improve large language models’ ability to impersonate real individuals in social media contexts. It also proposes an evaluation protocol that supplies LLM evaluators with reference information about the target individual. Experiments on a new dataset of social media replies show the method surpasses state‑of‑the‑art in‑context learning baselines, and the protocol correlates moderately with human judgments, while additional tests on fictional characters confirm broader applicability.
arXiv:2604. 24079v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) reveal inherent and distinctive personas through dialogue.
arXiv:2607. 26473v1 Announce Type: new Abstract: Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on explicit preference supervision such as pairwise comparisons or demographic attributes, limiting their applicability in natural interaction settings.