Introducing Falcon-H1-Arabic: Pushing the Boundaries of Arabic Language AI with Hybrid Architecture
Related stories
Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap
The paper surveys the state of Explainable AI (XAI) in Arabic NLP, highlighting three gaps: a method gap where Arabic XAI relies mainly on limited post‑hoc techniques; a task gap with most work focused on classification tasks and little on generation, retrieval, or dialogue; and a linguistic gap where explanations rarely address Arabic‑specific phenomena such as morphology, dialects, and diglossia. It proposes a taxonomy of tasks, methods, linguistic units, and evaluation practices, and outlines a research agenda for linguistically grounded Arabic XAI.
Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More
Jais 2: A Family of Arabic-Centric Open Large Language Models
arXiv:2608. 13580v1 Announce Type: cross Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report.
TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar)
The paper introduces CLASP‑Ar, a cloze‑style prompting method for Arabic stance detection that replaces complex multitask learning and ensembles with a single masked language modeling prompt. By combining the target, predicted sentiment, and text into one prompt and constraining the [MASK] prediction to a verbalizer‑defined label set, the approach aims to simplify the task while maintaining performance.
Nuha-Speech: Building General-Purpose Arabic Speech-LLMs
Nuha‑Speech is a new initiative aimed at creating general‑purpose Arabic speech‑large language models (speech‑LLMs). It includes the construction of a large Arabic Speech Question‑Answering corpus with over 1.5 million samples for instruction tuning, supervised fine‑tuning of Qwen‑Omni model variants at various scales, and a systematic evaluation framework with diverse tasks and tailored metrics. The project seeks to establish foundational infrastructure for Arabic speech‑LLMs amid limited Arabic speech resources.
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
Rosetta at AlexandriaX-2026: LoRA-Adapted NileChat for Context-Aware Dialectal Arabic Dialogue Translation
arXiv:2609.10395v1 Announce Type: new Abstract: This paper describes the Rosetta system for Subtask 1 (Context-Aware English-to-Dialectal Arabic Dialogue Translation) of the AlexandriaX shared task,...
Redteaming Leading Arabic LLMs with ASAS
arXiv:2608.21985v1 Announce Type: new Abstract: As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasingly critical...
Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge
The paper introduces Muslim, an Arabic voice AI platform that delivers grounded Islamic knowledge to users in real time. It combines a NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, and a self-hosted TTS system, supported by a deterministic multi-source retrieval layer across six Model Context Protocol servers. The authors release fine‑tuned Arabic Islamic model artifacts, describe an account‑based metering layer that prevents abuse, and present a three‑layer observability stack that monitors GPU‑bound agent hosts, reporting latency, accuracy, and engineering trade‑offs for production deployment.
BARRAC: Adaptation of an English Aspect-based Sentiment Analysis Approach for Classification Tasks in Arabic Dialects
The paper introduces BARRAC, a method that adapts an English aspect‑based sentiment analysis framework for Arabic dialect classification tasks. It replaces English consumer‑review attribute pools with Arabic linguistic markers for sentiment, sarcasm, and dialect identification, and swaps noisy self‑training for a two‑stage training process. Evaluated on five Arabic dialect datasets, BARRAC achieves a mean macro‑F1 of 63.93%, surpassing the best few‑label state‑of‑the‑art by 3% and outperforming GPT‑4o on four of the five tasks, while error analysis highlights remaining challenges.
Arabic Morphosyntactic Tagging and Dependency Parsing with Large Language Models
The paper evaluates large language models (LLMs) on Arabic morphosyntactic tagging and dependency parsing, a challenging task due to rich morphology and orthographic ambiguity. It compares zero‑shot prompting with retrieval‑based in‑context learning across pre‑tokenized, raw‑text, and cascaded settings, finding that relevant demonstrations significantly boost performance. The best LLMs nearly match supervised systems but need extensive annotated data for demonstrations and high computational resources. All code and data are publicly released.