Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance
Related stories
Falcon-Arabic: A Breakthrough in Arabic Language Models
Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs
The paper investigates how Arabic dialects are represented in large language models and whether they can be steered at inference time. By analyzing neuron-level sparsity and vector steering, the authors find that only a small fraction of neurons encode dialect-specific features, while distributed activation directions are more effective for steering. Vector steering can induce dialectal output from both dialectal and MSA prompts, whereas neuron steering works only when the prompt is already dialectal.
Some Dialects Are More Equal Than Others: Non-Prestigious Arabic Dialectal Bias in LLMs
arXiv:2609.23955v1 Announce Type: new Abstract: Previous work on Egyptian Arabic in NLP has focused largely on the prestigious Cairene Egyptian Arabic (CEA) dialect, resulting in a lack of representa...
Some Dialects Are More Equal Than Others: Non-Prestigious Arabic Dialectal Bias in LLMs
Previous work on Egyptian Arabic in NLP has focused largely on the prestigious Cairene Egyptian Arabic (CEA) dialect, resulting in a lack of representation for the less prestigious Sa'idi Egyptian Ara...
Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
The paper investigates cultural biases in large language models (LLMs) by introducing the Culture-Related Open Questions (CROQ) dataset, which contains 24‑language questions about generic culture. Experiments reveal that LLMs disproportionately favor Japan in their responses, especially when prompted in high‑resource languages, while low‑resource languages tend to highlight countries where the language is official. The study also finds that these biases emerge after supervised fine‑tuning rather than during pre‑training.
🇵🇭 FilBench - Can LLMs Understand and Generate Filipino?
Evaluating Cultural Awareness of LLMs for Haitian Creole
The paper presents the first systematic evaluation of cultural awareness in large language models (LLMs) for Haitian Creole, a low‑resource language. Using a benchmark of culturally salient prompts curated by native speakers, the study assesses four dimensions—specificity, bias, diversity, and variation—in a text‑infilling setting. Results show a clear gap between Haitian Creole and higher‑resource French, with Haitian performance more uneven and more affected by French linguistic interference; story generation reveals stereotypical portrayals of Haitian characters even in positive contexts.
Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?
arXiv:2609. 28245v1 Announce Type: cross Abstract: Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored.
EDRAC: Benchmarking Arabic Dialect Reading Comprehension
EDRAC is the first large‑scale benchmark for dialectal Arabic machine reading comprehension and generative question answering, covering five major dialects—Egyptian, Moroccan, Emirati, Syrian, and Saudi. It contains 499 passages from naturally spoken interactions and 4,977 QA pairs produced via a human–LLM collaborative pipeline. The benchmark evaluates Arabic‑centric and multilingual large language models, revealing gaps between semantic answer quality and dialectal fidelity and underscoring limitations of current evaluation metrics for dialectal Arabic generation.
How LLMs Might Think
arXiv:2604. 09674v2 Announce Type: replace Abstract: Do large language models (LLMs) think?