🇵🇭 FilBench - Can LLMs Understand and Generate Filipino?
Related stories
Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance
Open-source LLMs as LangChain Agents
Alyah ⭐️: Toward Robust Evaluation of Emirati Dialect Capabilities in Arabic LLMs
CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance
arXiv:2608.21462v1 Announce Type: cross Abstract: Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alp...
Llama 3.1 - 405B, 70B & 8B with multilinguality and long context
Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
The paper investigates cultural biases in large language models (LLMs) by introducing the Culture-Related Open Questions (CROQ) dataset, which contains 24‑language questions about generic culture. Experiments reveal that LLMs disproportionately favor Japan in their responses, especially when prompted in high‑resource languages, while low‑resource languages tend to highlight countries where the language is official. The study also finds that these biases emerge after supervised fine‑tuning rather than during pre‑training.
Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms
The paper introduces a generation benchmark for culturally specific kinship terms, evaluating five open‑weight large language models (LLMs) on Hindi, Tamil, and Korean. Unlike prior multiple‑choice tests that treat kinship understanding as a recognition task, the study prompts LLMs to generate terms across two communicative tasks and compares results to a matched option‑supported baseline. Findings show that while models like GPT‑OSS120B and Llama‑3.370B can select correct terms in over 90% and 78% of cases respectively, they produce the correct term only 36% and 24% of the time, indicating a significant evaluation‑format gap and highlighting the difficulty of culturally specific kinship generation even when relationships are explicitly stated.
How LLMs Might Think
arXiv:2604. 09674v2 Announce Type: replace Abstract: Do large language models (LLMs) think?
Letting Large Models Debate: The First Multilingual LLM Debate Competition
Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms
Current literature evaluates large language models (LLMs) on multilingual kinship understanding using multiple choice benchmarks, treating it as a recognition problem. We instead prompt five open weig...
Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?
arXiv:2607. 23440v1 Announce Type: cross Abstract: In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination.