Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
The paper investigates how Arabic dialects are represented in large language models and whether they can be steered at inference time. By analyzing neuron-level sparsity and vector steering, the authors find that only a small fraction of neurons encode dialect-specific features, while distributed activation directions are more effective for steering. Vector steering can induce dialectal output from both dialectal and MSA prompts, whereas neuron steering works only when the prompt is already dialectal.
arXiv:2609.23955v1 Announce Type: new Abstract: Previous work on Egyptian Arabic in NLP has focused largely on the prestigious Cairene Egyptian Arabic (CEA) dialect, resulting in a lack of representa...
Previous work on Egyptian Arabic in NLP has focused largely on the prestigious Cairene Egyptian Arabic (CEA) dialect, resulting in a lack of representation for the less prestigious Sa'idi Egyptian Ara...
The paper investigates cultural biases in large language models (LLMs) by introducing the Culture-Related Open Questions (CROQ) dataset, which contains 24‑language questions about generic culture. Experiments reveal that LLMs disproportionately favor Japan in their responses, especially when prompted in high‑resource languages, while low‑resource languages tend to highlight countries where the language is official. The study also finds that these biases emerge after supervised fine‑tuning rather than during pre‑training.