arXiv Machine Learning By E. Cho Smith, Samuel Ho, Dawn Laux

A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models

Read the original on arXiv Machine Learning →

The paper evaluates five large language models as zero‑shot annotators of four social constructs—self‑esteem, self‑control, seeking belonging, and seeking recognition—in English song lyrics. It examines repeated‑measurement reliability, cross‑model convergence, and the transferability of consensus labels to supervised classification. Results show varying reliability across constructs, with self‑esteem being most stable and seeking recognition least stable, and indicate that consensus labels contain learnable signal for downstream tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 3

Probing Cultural Signals in Large Language Models through Author Profiling

The study investigates cultural biases in large language models (LLMs) by testing their ability to perform author profiling—inferring singers’ gender and ethnicity—from song lyrics in a zero‑shot setting. Evaluating over 10,000 lyrics across several open‑source models, the authors find that most LLMs default toward North American ethnicity, while DeepSeek‑1.5B leans toward Asian ethnicity, and that Ministral‑8B exhibits the strongest ethnicity bias whereas Gemma‑12B is the most balanced. The paper introduces two fairness metrics, Modality Accuracy Divergence (MAD) and Recall Divergence (RD), to quantify these disparities and provides code and results publicly on GitHub and HuggingFace.

By Valentin Lafargue, Ariel Guerra-Adames, Emmanuelle Claeys, Elouan Vuichard, Jean-Michel Loubes
arXiv AI
Sep 3

Accurate in space, unreliable in time: how LLMs represent national cultural change

The paper investigates how large language models (LLMs) represent national cultural change over time, using more than two decades of World Values Survey data and the Inglehart‑Welzel cultural map. It finds that while LLMs generally place countries near their most recent surveyed positions, their representations lag behind current data, under‑capture the magnitude of change, introduce spurious movements, and rarely reproduce trajectory reversals. These temporal inaccuracies reveal a flattening effect that limits the models’ cultural awareness and raises concerns for evaluation, representational harms, and governance of culturally aware AI systems.

By Yalda Daryani, Miranda Bogen, Madeleine I. G. Daepp