Alyah ⭐️: Toward Robust Evaluation of Emirati Dialect Capabilities in Arabic LLMs
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
Previous work on Egyptian Arabic in NLP has focused largely on the prestigious Cairene Egyptian Arabic (CEA) dialect, resulting in a lack of representation for the less prestigious Sa'idi Egyptian Ara...
arXiv:2605.16364v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) voice assistants are commonly built as cascaded Automatic Speech recognition (ASR) to LLM systems, where recogni...
arXiv:2609.23955v1 Announce Type: new Abstract: Previous work on Egyptian Arabic in NLP has focused largely on the prestigious Cairene Egyptian Arabic (CEA) dialect, resulting in a lack of representa...
arXiv:2608.21950v1 Announce Type: cross Abstract: Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resource...
arXiv:2608.01291v2 Announce Type: replace Abstract: We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syri...
The paper investigates how Arabic dialects are represented in large language models and whether they can be steered at inference time. By analyzing neuron-level sparsity and vector steering, the authors find that only a small fraction of neurons encode dialect-specific features, while distributed activation directions are more effective for steering. Vector steering can induce dialectal output from both dialectal and MSA prompts, whereas neuron steering works only when the prompt is already dialectal.