Alyah โญ๏ธ: Toward Robust Evaluation of Emirati Dialect Capabilities in Arabic LLMs
Related stories
๐ 3LM: A Benchmark for Arabic LLMs in STEM and Code
Introducing the Open Arabic LLM Leaderboard
The Open Arabic LLM Leaderboard 2
Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction
arXiv:2607. 19751v1 Announce Type: cross Abstract: We present a regression-based approach to Arabic dialect geolocation that models dialectal variation as a continuous geographic space rather than discrete categories.
Introducing Falcon-H1-Arabic: Pushing the Boundaries of Arabic Language AI with Hybrid Architecture
Jais 2: A Family of Arabic-Centric Open Large Language Models
arXiv:2608. 13580v1 Announce Type: cross Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report.
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs
arXiv:2601. 12494v3 Announce Type: replace-cross Abstract: Audio large language models (LLMs) enable unified speech understanding and generation, but adapting them to linguistically complex and dialect-rich settings such as Arabic-English remains challenging.
DialectLLM: A Dialect-Aware Dialog[ue] Generation Framework Beyond Standard American English
arXiv:2601. 22888v4 Announce Type: replace-cross Abstract: More than 80% of the 1.
Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness
Dialectal variation remains a major challenge for multilingual language models. Perturbation-based continued pre-training (CPT) has emerged as a promising approach to improving robustness, yet existing work largely evaluates individual perturbation strategies in isolation and provides limited insight into why they work.
UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding
arXiv:2606. 07167v1 Announce Type: cross Abstract: Meaningful multilingual evaluation must test models in the target language and educational context.
PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors
arXiv:2606. 20137v1 Announce Type: cross Abstract: Existing mean opinion score (MOS) prediction models typically predict utterance-level naturalness MOS and can be insensitive to localized pitch-accent errors.