arXiv Computation and Language

MAPLE: Metadata Conditioned LLM Pretraining for Locale-Aware Question Answering

The paper introduces MAPLE, a family of decoder‑only language models pretrained with document‑level geographic metadata such as source URL, country, and continent. MAPLE is evaluated on a new benchmark, LocalNewsQA, which tests whether models can switch answers when the locale changes. Experiments show that, with inference‑time metadata fixed, MAPLE outperforms metadata‑free controls in both answer switching and accuracy on locale‑dependent questions, and these gains grow with model size.

arXiv AI
Jun 8

The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

arXiv:2606. 07422v1 Announce Type: cross Abstract: Large language models are increasingly used to answer culturally grounded questions across languages, yet it remains unclear whether local cultural knowledge is better accessed through English or the local language.

By Yang Zhang, Xiao Fei, Amr Mohamed, Sarah Almeida Carneiro, Mersin Konomi, Mingmeng Geng, Ahmed Asaad, Guokan Shang, Michalis Vazirgiannis
arXiv AI
Jul 15

Scaling Point-in-Time Language Models

arXiv:2607. 11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences.

By Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu
arXiv Machine Learning
Aug 18

L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

arXiv:2608. 15535v1 Announce Type: cross Abstract: We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs).

By Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi
arXiv Machine Learning
3d ago

The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline

The study investigates why language models exhibit systematic performance gaps across English dialects, a phenomenon termed the "dialect tax." Using parallel dialect corpora that preserve meaning while altering surface form, the authors confirm that models treat Standard American English and dialectal texts as semantically equivalent, yet find representational disparities that persist through tokenization, pre‑training, post‑training, and inference. Even a character‑level tokenizer does not eliminate input/output asymmetries or accuracy gaps, and dialect pairs produce more divergent gradient updates than unrelated Standard texts, indicating that dialectal content is harder for models to learn.

By Elle