arXiv Computation and Language

From annotation to reasoning: Culture in language models

The paper proposes a new way to evaluate language models on cultural understanding by focusing on interpretive depth rather than just factual recall. It argues that literary interpretation, where scholars can disagree yet still assess the quality of evidence, provides a useful framework for testing how models handle cultural references, reuse, and transformation across texts. The authors suggest linking evidence-centered benchmarks, preserving scholarly disagreement, and conducting model-development experiments on literary data, with Danish literature as a starting point for broader applications.

arXiv Computation and Language
Sep 10

AI translation of literary texts is "fine", but readers still prefer human translations

arXiv:2606.26040v2 Announce Type: replace Abstract: AI translation of literary works is increasingly common. While the content may be rendered adequately, we do not know enough about how readers expe...

By Yves Ferstler, Adam Podoxin, Ty Brassington, Ga\"elle Laperri\`ere, Roman Grundkiewicz, Marie-Jean Meurs, Maite Taboada, Marzena Karpinska
arXiv AI
Sep 17

Knowledge-Graph Based Augmentation versus Retrieval Augmented Generation for Cultural-Related Question Answering

The paper compares Knowledge-Graph Based Augmentation (Graph-RAG) with Retrieval-Augmented Generation (RAG) for answering culturally specific questions. Using the LatamQA dataset, Graph-RAG, built automatically from Wikipedia via KGGen, matches RAG performance and reduces the base LLM’s error by 72% with a standard KG and 78% with a benchmark-aware variant. The approach also transfers zero‑shot to Portuguese, showing multilingual applicability.

By Pablo Poulenard, Yannis Karmim, Valentin Barri\`ere
arXiv Machine Learning
Aug 5

VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP

arXiv:2608. 03095v1 Announce Type: cross Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese.

By Tu Tran Do, Nhat Ngoc Nguyen, Khanh-Tung Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang
arXiv Computation and Language
Sep 10

SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection

SWORD is a new benchmark that tests large language models’ ability to reject factually incorrect statements across eight major languages by distorting Wikidata triples. The benchmark reveals that models often perform better on semantically plausible distortions than on random ones, indicating a reliance on distributional familiarity rather than true factual verification. It also shows significant performance drops for East Asian languages, with gaps up to 28 percentage points, highlighting asymmetric multilingual factual reasoning capabilities.

By Sanghyeok Park, Minji Kang, Hosung Kwak, Jinhyuk Yun