The study investigates cultural biases in large language models (LLMs) by testing their ability to perform author profiling—inferring singers’ gender and ethnicity—from song lyrics in a zero‑shot setting. Evaluating over 10,000 lyrics across several open‑source models, the authors find that most LLMs default toward North American ethnicity, while DeepSeek‑1.5B leans toward Asian ethnicity, and that Ministral‑8B exhibits the strongest ethnicity bias whereas Gemma‑12B is the most balanced. The paper introduces two fairness metrics, Modality Accuracy Divergence (MAD) and Recall Divergence (RD), to quantify these disparities and provides code and results publicly on GitHub and HuggingFace.
By Valentin Lafargue, Ariel Guerra-Adames, Emmanuelle Claeys, Elouan Vuichard, Jean-Michel Loubes
arXiv:2606. 29273v1 Announce Type: cross Abstract: Emotion recognition of song lyrics is a challenging task since lyrics may not necessarily align with the overall emotion of a song.
By Rashini Liyanarachchi, Frank Tran, Md Mahmudul Hasan, Aditya Joshi, Erik Meijering
arXiv:2510. 08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts.
By Nikhil Reddy Varimalla, Yunfei Xu, Meng Fan Wang, Arkadiy Saakyan, Smaranda Muresan
arXiv:2608.28986v1 Announce Type: new
Abstract: LLMs often struggle with modern Korean poetry, producing outputs that resemble "line-broken prose." We address two coupled tasks: detecting whether a K...
By Keunhyeung Park, Seunguk Yu, YoungBin Kim
arXiv:2609.00565v1 Announce Type: cross
Abstract: Cultural fine-tuning has become the de facto paradigm for building culture-aware large language models (LLMs), yet existing optimization exclusively...
By Jingshen Zhang, Shaoyang Xu, Wenxuan Zhang
The paper investigates how large language models (LLMs) represent national cultural change over time, using more than two decades of World Values Survey data and the Inglehart‑Welzel cultural map. It finds that while LLMs generally place countries near their most recent surveyed positions, their representations lag behind current data, under‑capture the magnitude of change, introduce spurious movements, and rarely reproduce trajectory reversals. These temporal inaccuracies reveal a flattening effect that limits the models’ cultural awareness and raises concerns for evaluation, representational harms, and governance of culturally aware AI systems.
By Yalda Daryani, Miranda Bogen, Madeleine I. G. Daepp