arXiv Computation and Language By Haneul Yoo, Won Ik Cho, Geunhye Kim, Jiyoon Han

From National Curricula to Cultural Awareness: Constructing Open-Ended Culture-Specific Question Answering Dataset

Read the original on arXiv Computation and Language →

The paper introduces CuCu, a multi‑agent LLM framework that converts national social studies curricula into open‑ended, culture‑specific question‑answer pairs for fine‑tuning language models. Using the Korean curriculum, the authors build KCaQA, a dataset of 34.1k QA pairs that cover culture‑specific topics and ground responses in local sociocultural contexts. Experiments show that fine‑tuning with KCaQA improves the model’s cultural alignment and relevance to Korean society.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
4d ago

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

arXiv:2608.28405v1 Announce Type: new Abstract: Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common...

By Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan, Bich Ngoc Doan, Jia Wang Peh, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Zhengyuan Liu, JinYeong Bak, Wafi Shamdi, Soo Kai Chie, Liew Yu Siong, Aina Azyyati Binti Mohamad Rezal, Lew Yan Yan Vanessa, Huadan Wu, Dylan Raharja, Nadya Yuki Wangsajaya, Akane Fukushige, Kazushi Kato, Koji Inoue, Tatsuya Kawahara, Jaehyung Seo, Dongjun Kim, Seungyoon Lee, Zi Haur Pang, Rui Yang Tan, Charibeth Ko Cheng, Maria Regina Justina Estuar, Jann Railey Montalan, Pham Minh Duc, Roy Ka-Wei Lee
arXiv AI
Aug 28

Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content

The paper introduces MentorQA, a multilingual dataset and evaluation framework for mentorship-oriented question answering derived from long‑form videos. It contains nearly 9,000 QA pairs across four languages and defines evaluation dimensions such as clarity, alignment, and learning value that extend beyond factual accuracy. Experiments show that Multi‑Agent QA pipelines outperform other architectures, especially on complex topics and low‑resource languages, while automated LLM‑based evaluation shows variable alignment with human judgments.

By Parth Bhalerao, Diola Dsouza, Ruiwen Guan, Oana Ignat
arXiv Computation and Language
4d ago

Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs

The paper investigates cultural biases in large language models (LLMs) by introducing the Culture-Related Open Questions (CROQ) dataset, which contains 24‑language questions about generic culture. Experiments reveal that LLMs disproportionately favor Japan in their responses, especially when prompted in high‑resource languages, while low‑resource languages tend to highlight countries where the language is official. The study also finds that these biases emerge after supervised fine‑tuning rather than during pre‑training.

By Joseba Fernandez de Landa, Carla Perez-Almendros, Jose Camacho-Collados
arXiv Machine Learning
Aug 18

L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

arXiv:2608. 15535v1 Announce Type: cross Abstract: We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs).

By Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi
arXiv AI
Aug 21

DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

arXiv:2509. 08022v3 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks often overlook cultural and demographic variation.

By Yao Liang, Dongcheng Zhao, Feifei Zhao, Guobin Shen, Yuwei Wang, Dongqi Liang, Yi Zeng