arXiv Computation and Language By Mohammed Elkholy, Khalid N. Elmadani, Nizar Habash, Bashar Alhafni

Arabic Sentence Segmentation Across Genres and Punctuation Conditions

Read the original on arXiv Computation and Language →

The paper introduces AraSEG, a new Arabic sentence segmentation corpus covering eight genres and diverse punctuation and document structures. Experiments using AraSEG evaluate large language models, lightweight encoders, and dependency parser-based models, revealing that lightweight encoders and parser-based models outperform LLMs under the most challenging conditions. The study also shows that increasing training data size and genre diversity eventually saturates performance, that cross‑genre generalization remains difficult, and that accurate sentence segmentation significantly improves downstream dependency parsing.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 4

Arabic Morphosyntactic Tagging and Dependency Parsing with Large Language Models

The paper evaluates large language models (LLMs) on Arabic morphosyntactic tagging and dependency parsing, a challenging task due to rich morphology and orthographic ambiguity. It compares zero‑shot prompting with retrieval‑based in‑context learning across pre‑tokenized, raw‑text, and cascaded settings, finding that relevant demonstrations significantly boost performance. The best LLMs nearly match supervised systems but need extensive annotated data for demonstrations and high computational resources. All code and data are publicly released.

By Mohamed Adel, Bashar Alhafni, Nizar Habash
arXiv Computation and Language
2d ago

Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

Nuha‑Speech is a new initiative aimed at creating general‑purpose Arabic speech‑large language models (speech‑LLMs). It includes the construction of a large Arabic Speech Question‑Answering corpus with over 1.5 million samples for instruction tuning, supervised fine‑tuning of Qwen‑Omni model variants at various scales, and a systematic evaluation framework with diverse tasks and tailored metrics. The project seeks to establish foundational infrastructure for Arabic speech‑LLMs amid limited Arabic speech resources.

By Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi
arXiv AI
Jun 12

AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages

arXiv:2606. 12708v1 Announce Type: cross Abstract: Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP.

By Happy Buzaaba, Cheikh Mouhamadou Bamba Dione, David Ifeoluwa Adelani, Sylvain Kahane, Kim Gerdes, Bruno Guillaume, Kevin Guan, Aremu Anuoluwapo, Naome A. Etori, Shamsuddeen Hassan Muhammad, Utitofon Inyang, Peter Nabende, David Sabiiti Bamutura, Andiswa Bukula, Chinedu Uchechukwu, Rooweither Mabuya, Idris Akinade, Christiane Fellbaum
Hugging Face Trending Papers
Jun 23

CANDLE: Character-level Arabic Noise Deduplication using Lightweight Encoder

Handling repeated characters in text can be tricky, since they can represent either the correct spelling of a word or informal character elongation often seen in social media posts. We present CANDLE, a lightweight system for character-level Arabic noise deduplication that addresses this challenge without relying on handcrafted rules, dictionaries, or morphological analyzers.

arXiv AI
Aug 17

Jais 2: A Family of Arabic-Centric Open Large Language Models

arXiv:2608. 13580v1 Announce Type: cross Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report.

By Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim, Mostafa Awad, Abdelrahman Sadallah, Gurpreet Gosal, Gokulakrishnan Ramakrishnan, Sarath Chandran, Biswajit Mishra, Rituraj Joshi, Ahmed Frikha, Etienne Goffinet, Abhishek Maiti, Ali El Filali, Sarah AlBarri, Samujjwal Ghosh, Rahul Pal, Parvez Mullah, Awantika Shukla, Sajid siddiki, Samta Kamboj, Onkar Pandit, Sunil Kumar Sahu, AbdelRahman Elbadawy, Amr Mohamed, Ahmad Chamma, Evan Dufraisse, Abdelaziz Bounhar, Dani Bouch, Hadi Abdine, Guokan Shang, Fajri Koto, Yuxia Wang, Zhuohan Xie, Ali Mekky, Rania Elbadry, Sarfraz Ahmad, Momina Ahsan, Omar El Herraoui, Daniil Orel, Hasan Iqbal, Kareem Elzeky, Mervat Abassy, Kareem Elozeiri, Saadeldine Eletter, Farah Atif, Nurdaulet Mukhituly, Haonan Li, Xudong Han, Aaryamonvikram Singh, Zainul Abedien Ahmed Quraishi, Neha Sengupta, Larry Murray, Avraham Sheinin, Joel Hestness, Natalia Vassilieva, Hector Xuguang Ren, Zhengzhong Liu, Michalis Vazirgiannis, Preslav Nakov