Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is...
The paper presents a knowledge graph built from Harvard Dataverse’s public data, linking 102,650 datasets to 215,985 nodes and 528,003 edges that include keywords, publications, subjects, journals, and locations. About 43,991 datasets contain geospatial metadata, and 7,654 are identified as policy‑relevant, with elections and legislatures forming the largest cluster. The authors highlight the challenge of place resolution—disconnected nodes representing the same location—and propose the graph as a testbed for AI‑driven metadata enrichment and entity resolution, noting a bias toward American city‑level data.
By Danny EBanks, Devika Jain
arXiv:2608. 07411v1 Announce Type: new Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities.
By Rodrigo Ferreira Rodrigues, Karim Radouane, Jose G Moreno, Lynda Tamine
The paper introduces MAPLE, a family of decoder‑only language models pretrained with document‑level geographic metadata such as source URL, country, and continent. MAPLE is evaluated on a new benchmark, LocalNewsQA, which tests whether models can switch answers when the locale changes. Experiments show that, with inference‑time metadata fixed, MAPLE outperforms metadata‑free controls in both answer switching and accuracy on locale‑dependent questions, and these gains grow with model size.
By Anjishnu Mukherjee, Ziwei Zhu, Antonios Anastasopoulos
arXiv:2607. 06482v1 Announce Type: cross Abstract: Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings.
By So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen
arXiv:2405.06818v2 Announce Type: replace
Abstract: Natural Language Processing (NLP) for Ghana's 73 living indigenous languages remains deeply fragmented, under-resourced, and heavily skewed toward...
By Sheriff Issaka, Erick Rosas Gonzalez, Colene Agbo, Evans Kofi Agyei, Shruti Tyagi, John Emeka Eze, Enock Appiah Tieku, Junlin Fang, Thanh Do Nguyen, Juliet Arthur, Zhaoyi Zhang, Mihir Heda, Keyi Wang, Yinka Ajibola, Rebecca Akpanglo-Nartey, Frank Lawrence Nii Adoquaye Acquaye, Dennis Owusu, Jerry John Kponyo, Stephen Moore, Isaac Wiafe, Sean Du
We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language models (LLMs) to perform fine-grained reasoning over geographically grounded and regionally differentiated knowledge within France. While prior benchmarks focus on national-level cultural understanding, they largely overlook intra-country variation and the need to distinguish between closely related regional contexts.
arXiv:2602.14488v3 Announce Type: replace-cross
Abstract: IR in low-resource languages remains limited by the scarcity of high-quality, task-specific annotated datasets. Manual annotation is expensiv...
By Md. Najib Hasan, Mst. Jannatun Ferdous Rain, Fyad Mohammed, Nazmul Siddique
arXiv:2606. 02255v1 Announce Type: cross Abstract: Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and how the annotation process was controlled.
By Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli, Yanran Chen, Christian Greisinger, Lotta Kiefer, Christoph Leiter, Subhadeep Roy, Tewodros Achamaleh, Muhammad Arslan Manzoor, Sebastian Pohl, Yufang Hou, Steffen Eger
arXiv:2609.22494v1 Announce Type: new
Abstract: In recent years, there has been a surge of interest in Cultural NLP, with substantial efforts to create globally inclusive NLP systems. The rapid growt...
By Tania Chakraborty, Eylon Caplan, Zhaoqing Wu, Kevin Cushing, Han Qin, Shreya Havaldar, Dan Goldwasser
arXiv:2606.18389v2 Announce Type: replace
Abstract: Large language models (LLMs) have become an effective tool for synthetic data generation, including for low-resource languages, where generated dat...
By Jan Cegin, Daniil Gurgurov, Yusser Al Ghussin, Simon Ostermann
arXiv:2609.16592v1 Announce Type: new
Abstract: This paper presents an end-to-end approach for generating context-specific large language model (LLM) benchmark datasets by combining expert input with...
By Kimberly Le Truong, Nari Johnson, Anna Kawakami, Hoda Heidari