arXiv Machine Learning

A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models

The paper evaluates five large language models as zero‑shot annotators of four social constructs—self‑esteem, self‑control, seeking belonging, and seeking recognition—in English song lyrics. It examines repeated‑measurement reliability, cross‑model convergence, and the transferability of consensus labels to supervised classification. Results show varying reliability across constructs, with self‑esteem being most stable and seeking recognition least stable, and indicate that consensus labels contain learnable signal for downstream tasks.

arXiv Machine Learning
Sep 3

Probing Cultural Signals in Large Language Models through Author Profiling

The study investigates cultural biases in large language models (LLMs) by testing their ability to perform author profiling—inferring singers’ gender and ethnicity—from song lyrics in a zero‑shot setting. Evaluating over 10,000 lyrics across several open‑source models, the authors find that most LLMs default toward North American ethnicity, while DeepSeek‑1.5B leans toward Asian ethnicity, and that Ministral‑8B exhibits the strongest ethnicity bias whereas Gemma‑12B is the most balanced. The paper introduces two fairness metrics, Modality Accuracy Divergence (MAD) and Recall Divergence (RD), to quantify these disparities and provides code and results publicly on GitHub and HuggingFace.

By Valentin Lafargue, Ariel Guerra-Adames, Emmanuelle Claeys, Elouan Vuichard, Jean-Michel Loubes
arXiv AI
Sep 3

Accurate in space, unreliable in time: how LLMs represent national cultural change

The paper investigates how large language models (LLMs) represent national cultural change over time, using more than two decades of World Values Survey data and the Inglehart‑Welzel cultural map. It finds that while LLMs generally place countries near their most recent surveyed positions, their representations lag behind current data, under‑capture the magnitude of change, introduce spurious movements, and rarely reproduce trajectory reversals. These temporal inaccuracies reveal a flattening effect that limits the models’ cultural awareness and raises concerns for evaluation, representational harms, and governance of culturally aware AI systems.

By Yalda Daryani, Miranda Bogen, Madeleine I. G. Daepp
arXiv Computation and Language
3d ago

Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling

arXiv:2609.22934v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characteriz...

By Yu Sha, Junqi Tao, Dixin Zhou, Yansheng Tu, Mingyang Chen, Xiang Fan, Yang Liu, Mengquan Yang, Jie Lin, Jiahui Fu, Hua Zheng, Benwei Zhang, Zhou Kai