The study investigates cultural biases in large language models (LLMs) by testing their ability to perform author profiling—inferring singers’ gender and ethnicity—from song lyrics in a zero‑shot setting. Evaluating over 10,000 lyrics across several open‑source models, the authors find that most LLMs default toward North American ethnicity, while DeepSeek‑1.5B leans toward Asian ethnicity, and that Ministral‑8B exhibits the strongest ethnicity bias whereas Gemma‑12B is the most balanced. The paper introduces two fairness metrics, Modality Accuracy Divergence (MAD) and Recall Divergence (RD), to quantify these disparities and provides code and results publicly on GitHub and HuggingFace.
By Valentin Lafargue, Ariel Guerra-Adames, Emmanuelle Claeys, Elouan Vuichard, Jean-Michel Loubes
arXiv:2606. 29273v1 Announce Type: cross Abstract: Emotion recognition of song lyrics is a challenging task since lyrics may not necessarily align with the overall emotion of a song.
By Rashini Liyanarachchi, Frank Tran, Md Mahmudul Hasan, Aditya Joshi, Erik Meijering
arXiv:2510. 08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts.
By Nikhil Reddy Varimalla, Yunfei Xu, Meng Fan Wang, Arkadiy Saakyan, Smaranda Muresan
arXiv:2608.28986v1 Announce Type: new
Abstract: LLMs often struggle with modern Korean poetry, producing outputs that resemble "line-broken prose." We address two coupled tasks: detecting whether a K...
By Keunhyeung Park, Seunguk Yu, YoungBin Kim
arXiv:2609.00565v1 Announce Type: cross
Abstract: Cultural fine-tuning has become the de facto paradigm for building culture-aware large language models (LLMs), yet existing optimization exclusively...
By Jingshen Zhang, Shaoyang Xu, Wenxuan Zhang
The paper investigates how large language models (LLMs) represent national cultural change over time, using more than two decades of World Values Survey data and the Inglehart‑Welzel cultural map. It finds that while LLMs generally place countries near their most recent surveyed positions, their representations lag behind current data, under‑capture the magnitude of change, introduce spurious movements, and rarely reproduce trajectory reversals. These temporal inaccuracies reveal a flattening effect that limits the models’ cultural awareness and raises concerns for evaluation, representational harms, and governance of culturally aware AI systems.
By Yalda Daryani, Miranda Bogen, Madeleine I. G. Daepp
arXiv:2609.15511v1 Announce Type: cross
Abstract: This paper investigates the generation and human evaluation of Japanese haiku by contemporary Large Language Models (LLMs), focusing on authorship pe...
By Livia Oddi, Simone Scardapane, Toru Sugimoto, Donatella Genovese
arXiv:2609.22934v1 Announce Type: new
Abstract: Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characteriz...
By Yu Sha, Junqi Tao, Dixin Zhou, Yansheng Tu, Mingyang Chen, Xiang Fan, Yang Liu, Mengquan Yang, Jie Lin, Jiahui Fu, Hua Zheng, Benwei Zhang, Zhou Kai
arXiv:2609.00491v1 Announce Type: new
Abstract: Communicating across cultures is inherently challenging, especially through culturally dense and ambiguous formats like memes. While people expect larg...
By Hangxiao Zhu, Suliu Qin, Zhuoyan Li, Ming Jiang, Yu Zhang, Meng Xia
arXiv:2607. 20241v1 Announce Type: cross Abstract: Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms.
By Yiming Wang, Jiayuan Di
arXiv:2608.29266v1 Announce Type: cross
Abstract: Researchers increasingly treat LLM survey responses as a proxy for human cultural values. This includes projecting model outputs onto instruments lik...
By An Duy Nguyen, Muhammad Aurangzeb Ahmad
arXiv:2604. 16009v2 Announce Type: replace Abstract: Most large language model benchmarks evaluate final-answer quality but reveal little about how models revise beliefs under disagreement or conflicting evidence.
By Farhad Abtahi, Abdolamir Karbalaie, Eduardo Illueca-Fernandez, Fernando Seoane