AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2608.30107v1 Announce Type: cross Abstract: Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and i...
arXiv:2607. 06482v1 Announce Type: cross Abstract: Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings.
The paper presents a knowledge graph built from Harvard Dataverse’s public data, linking 102,650 datasets to 215,985 nodes and 528,003 edges that include keywords, publications, subjects, journals, and locations. About 43,991 datasets contain geospatial metadata, and 7,654 are identified as policy‑relevant, with elections and legislatures forming the largest cluster. The authors highlight the challenge of place resolution—disconnected nodes representing the same location—and propose the graph as a testbed for AI‑driven metadata enrichment and entity resolution, noting a bias toward American city‑level data.
The paper introduces MAPLE, a family of decoder‑only language models pretrained with document‑level geographic metadata such as source URL, country, and continent. MAPLE is evaluated on a new benchmark, LocalNewsQA, which tests whether models can switch answers when the locale changes. Experiments show that, with inference‑time metadata fixed, MAPLE outperforms metadata‑free controls in both answer switching and accuracy on locale‑dependent questions, and these gains grow with model size.
arXiv:2608. 07411v1 Announce Type: new Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities.
arXiv:2405.06818v2 Announce Type: replace Abstract: Natural Language Processing (NLP) for Ghana's 73 living indigenous languages remains deeply fragmented, under-resourced, and heavily skewed toward...