The paper explores how "culture" can be operationalised in Natural Language Processing (NLP) and what this reveals about the possibilities and limits of considering a plurality of cultural backgrounds in technological design. It proposes that cultural alignment cannot be achieved only by adding more examples of "other cultures", rather it requires plural epistemologies: allowing multiple, locally grounded ways of knowing.
QQ is a language metadata toolkit designed for multilingual NLP research. It aggregates diverse language metadata into a graph of varieties, scripts, regions, identifiers, names, and relations, and offers access via a Python API, CLI, and browser explorer. The toolkit enables normalization of identifiers, metadata retrieval, relation traversal, and discovery of external resources containing a language, and is demonstrated through audits of the HuggingFace Hub, linking resources with different identifier systems, and generating reproducible language-reporting tables.
By Wessel Poelman, Yiyi Chen, Miryam de Lhoneux
arXiv:2405.06818v2 Announce Type: replace
Abstract: Natural Language Processing (NLP) for Ghana's 73 living indigenous languages remains deeply fragmented, under-resourced, and heavily skewed toward...
By Sheriff Issaka, Erick Rosas Gonzalez, Colene Agbo, Evans Kofi Agyei, Shruti Tyagi, John Emeka Eze, Enock Appiah Tieku, Junlin Fang, Thanh Do Nguyen, Juliet Arthur, Zhaoyi Zhang, Mihir Heda, Keyi Wang, Yinka Ajibola, Rebecca Akpanglo-Nartey, Frank Lawrence Nii Adoquaye Acquaye, Dennis Owusu, Jerry John Kponyo, Stephen Moore, Isaac Wiafe, Sean Du
arXiv:2608.30107v1 Announce Type: cross
Abstract: Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and i...
By Joan Nwatu, Tsedeniya Solomon Amare, Longju Bai, Bontu Fufa Balcha, Zayd Bashir, Angana Borah, Zara Burzo, Yubin Choi, Naihao Deng, Samika Gupta, Michel Faloughi, Claude Kwizera, Ziqiao Ma, Cynthia Yacel Fuertes Panizo, Ellie Seehorn, Hui Shen, Jiayi Tang, Zesen Zhao, Boyuan Zheng, Rada Mihalcea
The paper "Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems" investigates how commercial text‑to‑image models silently modify user prompts before generating images, a step that is often hidden from users. Using the multilingual benchmark WORLDVIEW, the authors audit the revision layer in DALL‑E‑3, Imagen‑4, and GPT‑Image‑1.5, finding that non‑Western and non‑Anglophone contexts are disproportionately marked, reduced to narrow vocabularies, and stereotyped. The study demonstrates that the revision layer itself is a previously undocumented causal source of cultural stereotyping, underscoring the need to audit deployed systems rather than just the underlying models.
By Aleksandra Urman, Elsa Lichtenegger, Salima Jaoua, Azza Bouleimen, Robin Forsberg, Corinna Hertweck, Stefania Ionescu, Nicol\`o Pagan, Ancsa Hannak, Joachim Baumann
Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is...
arXiv:2608. 11002v1 Announce Type: cross Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years.
By Sicheng Zhang, Zhonghao Yan, Binzhu Xie, Shi Qiu, Muzammal Naseer, Naveed Akhtar, Mubarak Shah
arXiv:2607. 20241v1 Announce Type: cross Abstract: Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms.
By Yiming Wang, Jiayuan Di
arXiv:2607. 06544v1 Announce Type: new Abstract: As Artificial Intelligence (AI) makes inroads into different parts of the Indian subcontinent, there is significant interest in studying how AI impacts the linguistic and cultural foundations of this civilization.
By Aparna Madva, Sharath Srivatsa, Srinath Srinivasa, Tulika Saha
The paper introduces data stories as narrative documents that combine explanatory text, images, and executable SPARQL queries with visualized results to make cultural‑heritage knowledge graphs more accessible. It describes how these stories guide users through unfamiliar graphs, create reproducible narratives, and uncover hidden data‑quality issues. The authors present LODEON, an authoring platform with Sparnatural and AI‑assisted tools, and report early positive feedback from seminars and workshops.
By Tabea Tietz, Torsten Schrade, Etienne Posthumus, Linnaea S\"ohn, Jonatan Jalle Steller, J\"org Waitelonis, Harald Sack
Building language technologies and conducting NLP research for low-resource languages---particularly when led by native speakers or involving participatory research practices---are often framed as mea...
arXiv:2601. 14063v2 Announce Type: replace-cross Abstract: Cross-cultural competence in large language models (LLMs) requires understanding and adapting Culture-Specific Items (CSIs) across varying cultural contexts.
By Mohsinul Kabir, Tasnim Ahmed, Md Mezbaur Rahman, Shaoxiong Ji, Hassan Alhuzali, Yuechen Jiang, Jimin Huang, Sophia Ananiadou