CultureMINE: Datasets and Methods for Improving the Cultural Capabilities of NLP Systems
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper explores how "culture" can be operationalised in Natural Language Processing (NLP) and what this reveals about the possibilities and limits of considering a plurality of cultural backgrounds in technological design. It proposes that cultural alignment cannot be achieved only by adding more examples of "other cultures", rather it requires plural epistemologies: allowing multiple, locally grounded ways of knowing.
QQ is a language metadata toolkit designed for multilingual NLP research. It aggregates diverse language metadata into a graph of varieties, scripts, regions, identifiers, names, and relations, and offers access via a Python API, CLI, and browser explorer. The toolkit enables normalization of identifiers, metadata retrieval, relation traversal, and discovery of external resources containing a language, and is demonstrated through audits of the HuggingFace Hub, linking resources with different identifier systems, and generating reproducible language-reporting tables.
arXiv:2405.06818v2 Announce Type: replace Abstract: Natural Language Processing (NLP) for Ghana's 73 living indigenous languages remains deeply fragmented, under-resourced, and heavily skewed toward...
arXiv:2608.30107v1 Announce Type: cross Abstract: Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and i...
The paper "Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems" investigates how commercial text‑to‑image models silently modify user prompts before generating images, a step that is often hidden from users. Using the multilingual benchmark WORLDVIEW, the authors audit the revision layer in DALL‑E‑3, Imagen‑4, and GPT‑Image‑1.5, finding that non‑Western and non‑Anglophone contexts are disproportionately marked, reduced to narrow vocabularies, and stereotyped. The study demonstrates that the revision layer itself is a previously undocumented causal source of cultural stereotyping, underscoring the need to audit deployed systems rather than just the underlying models.
Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is...