Hugging Face Trending Papers

Beyond Good Intentions: When Does the Framing of Multilingual and Low-Resource NLP Research Become a Caricature?

arXiv Computation and Language
6d ago

A Systematic Review of NLP for Ghanaian Languages: Datasets, Models, and a Research Roadmap

arXiv:2405.06818v2 Announce Type: replace Abstract: Natural Language Processing (NLP) for Ghana's 73 living indigenous languages remains deeply fragmented, under-resourced, and heavily skewed toward...

By Sheriff Issaka, Erick Rosas Gonzalez, Colene Agbo, Evans Kofi Agyei, Shruti Tyagi, John Emeka Eze, Enock Appiah Tieku, Junlin Fang, Thanh Do Nguyen, Juliet Arthur, Zhaoyi Zhang, Mihir Heda, Keyi Wang, Yinka Ajibola, Rebecca Akpanglo-Nartey, Frank Lawrence Nii Adoquaye Acquaye, Dennis Owusu, Jerry John Kponyo, Stephen Moore, Isaac Wiafe, Sean Du
arXiv Computation and Language
Sep 4

Beyond Accuracy: Community Perspectives on Machine Translation

The paper examines how four stakeholder groups—AI developers, professional translators, language learners, and language service providers—discuss machine translation on social media. Using a dataset of 79,286 posts from Reddit, Facebook, Bluesky, and Mastodon (2019‑2025), the authors find frequent disagreements and strong conflicts over translation quality, efficiency, and reliability. These conflicts arise because AI communities view the issues as technical, while non‑AI users prioritize quality nuances, time savings, trust, and broader social concerns.

By Yujun Wang, Ehud Reiter, Shimei Pan, Steffen Eger, Wei Zhao
Hugging Face Trending Papers
Jun 8

Beyond Accuracy: Community Perspectives on Machine Translation

Despite remarkable progress in machine translation (MT), non-AI communities have raised growing concerns about MT systems, suggesting a noticeable gap between technical advancement and the needs of real-world users. For instance, while NLP researchers focus on benchmark performance, end users care about ethical concerns, trust, reliability, costs, and more.

arXiv Computation and Language
Sep 3

Are Non-English Papers Reviewed Fairly? Language-of-Study Bias in NLP Peer Reviews

The paper investigates language-of-study (LoS) bias in NLP peer reviews, defining and distinguishing negative and positive forms of bias. Using a new dataset, LOBSTER, and an LLM-based detection pipeline, the authors analyze 15,645 reviews and find that non‑English papers experience significantly higher bias rates, with negative bias outweighing positive bias. They further identify four subcategories of negative bias, noting that demanding unjustified cross‑lingual generalization is the most common.

By Ehsan Barkhordar, Abdulfattah Safa, Verena Blaschke, Erika Lombart, Marie-Catherine de Marneffe, G\"ozde G\"ul \c{S}ahin
Hugging Face Trending Papers
Jun 20

Plurification in/of language technology -- The integration of culture in next-generation AI

The paper explores how "culture" can be operationalised in Natural Language Processing (NLP) and what this reveals about the possibilities and limits of considering a plurality of cultural backgrounds in technological design. It proposes that cultural alignment cannot be achieved only by adding more examples of "other cultures", rather it requires plural epistemologies: allowing multiple, locally grounded ways of knowing.

arXiv AI
Jul 3

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

arXiv:2607. 02235v1 Announce Type: cross Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English.

By A. Seza Do\u{g}ru\"oz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani
arXiv AI
Jul 8

Rethinking Indic AI from a Lens of Cultural Heritage Preservation

arXiv:2607. 06544v1 Announce Type: new Abstract: As Artificial Intelligence (AI) makes inroads into different parts of the Indian subcontinent, there is significant interest in studying how AI impacts the linguistic and cultural foundations of this civilization.

By Aparna Madva, Sharath Srivatsa, Srinath Srinivasa, Tulika Saha
arXiv AI
Jun 12

Mod-Guide: An LLM-based Content Moderation Feedback System to Address Insensitive Speech toward Indigenous Ethnic and Religious Minority Communities

arXiv:2606. 13397v1 Announce Type: cross Abstract: Language operates as a mechanism of both marginalization and resistance, especially for minority communities navigating insensitive and harmful speech online.

By Dipto Das, Achhiya Sultana, Ankit Singh Chauhan, Saadia Binte Alam, Mohammad Shidujaman, Shion Guha, Sunandan Chakraborty, Syed Ishtiaque Ahmed
arXiv Computation and Language
Aug 25

Opportunities and Challenges of Natural Language Processing for Low-Resource Senegalese Languages in Social Science Research

arXiv:2601.09716v2 Announce Type: replace Abstract: Natural Language Processing (NLP) is rapidly transforming research methodologies across disciplines, yet African languages remain largely underrepr...

By Derguene Mbaye, Tatiana D. P. Mbengue, Madoune R. Seye, Moussa Diallo, Mamadou L. Ndiaye, Dimitri S. Adjanohoun, Cheikh S. Wade, Djiby Sow, Jean-Claude B. Munyaka, Jerome Chenal