The paper introduces a multilingual story moral generation task to evaluate cultural alignment in large language models. Using a dataset of human-written story morals from 14 language‑culture pairs, the authors compare model outputs to human interpretations through semantic similarity, a preference survey, and value categorization. They find that advanced models like GPT‑4o and Gemini produce morally similar and preferred responses but show less cross‑linguistic variation, focusing on a narrower set of shared values, indicating a limitation in capturing the diversity of human narrative understanding.
By Sophie Wu, Andrew Piper
The study introduces a World Values Survey–grounded simulation framework to test whether large language model agents can faithfully represent diverse human value systems. In about 4,000 conversations with 1,200 personas across three models, more than half of the agents failed to express their assigned value profiles from the start, and only 2–7% drifted over time. The results show systematic deviations from the intended value distributions and reveal that simulated dialogues differ from human discussions in their balance of stylistic consistency and semantic diversity.
By Farah Atif, Sougata Saha, Monojit Choudhury
arXiv:2607. 24782v1 Announce Type: new Abstract: LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast how people would answer value-laden questions.
By James Wedgwood, Pratiksha Thaker, Neil Kale, Virginia Smith
PERSONAWEAVER is a new approach to procedural character generation that separates world building from behavioral specification, using manually curated banks of moral positions and conversational reactions to diversify character behavior. By applying this method across ten realistic and fantastical settings and three large language models, the system produces broader moral and interactional response distributions, varied interpersonal language, response length, sentiment, and less archetypal world attribute combinations compared to prior work.
By Maan Qraitem, Kate Saenko, Bryan A. Plummer
arXiv:2604. 09945v2 Announce Type: replace-cross Abstract: The rapid adoption of large vision-language models (LVLMs) in recent years has been accompanied by growing fairness concerns due to their propensity to reinforce harmful societal stereotypes.
By Phillip Howard, Xin Su, Kathleen C. Fraser
arXiv:2605. 30036v2 Announce Type: replace Abstract: Large Language Models (LLMs) demonstrate a remarkable capacity to adopt different personas and roles; however, it remains unclear whether they can manifest behavior that adheres to a coherent, human-like value structure.
By Asaf Yehudai, Naama Rozen, Ariel Gera