arXiv:2608.23421v1 Announce Type: new
Abstract: Natural Language Processing (NLP) has grown rapidly over the past decade, driven by digital transformation in the Arab world, social media, and large l...
By Mullosharaf K. Arabov
The paper surveys the state of Explainable AI (XAI) in Arabic NLP, highlighting three gaps: a method gap where Arabic XAI relies mainly on limited post‑hoc techniques; a task gap with most work focused on classification tasks and little on generation, retrieval, or dialogue; and a linguistic gap where explanations rarely address Arabic‑specific phenomena such as morphology, dialects, and diglossia. It proposes a taxonomy of tasks, methods, linguistic units, and evaluation practices, and outlines a research agenda for linguistically grounded Arabic XAI.
By Salima Lamsiyah, Ruslan Mitkov
Russian and Arabic are among the major languages of scientific communication. Language barriers impede the exchange of research results between these communities, which affects international collaboration and the progress of sustainability-related research.
AraDetox is a newly released multi-dialect Arabic detoxification dataset containing 10,500 harmful social‑media posts and 84,000 detoxified rewrites generated by GPT‑5 and Gemini 2.5 Flash across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic. Human evaluation and automatic analyses confirm that the rewrites effectively remove harmful language while preserving meaning, lexical change, and dialectal style. The dataset is publicly available to support future research in Arabic detoxification, safe text generation, and multi‑dialect NLP.
By Mo El-Haj
arXiv:2607. 20056v1 Announce Type: cross Abstract: Aspect-based sentiment analysis (ABSA) in Arabic must recover both explicitly stated aspects and implicit aspects that are never named in the text.
By Lujain A. Alawwad
The article reviews how transformer models, retrieval‑grounded pipelines, and large language models are reshaping hadith computational science. It critically evaluates existing literature, highlighting uneven progress: expanded data resources and mature segmentation tasks, yet persistent issues such as narrow corpora, weak benchmark comparability, and limited reproducibility. The authors argue that hadith computation should be viewed as an evidence‑infrastructure problem requiring knowledge integration, provenance, and expert supervision, and they propose a research agenda to strengthen the field’s methodological rigor.
By Md. Ashraful Haque (Greentech Apps Foundation, United Kingdom), Riasat Islam (Greentech Apps Foundation, United Kingdom, Queen Mary University of London, London, United Kingdom)
arXiv:2606.22342v2 Announce Type: replace
Abstract: How does research evolve, and can we trace it at the level of individual claims? Scientific progress is not simply a uniform accumulation of facts....
By Abdul Muntakim, Md Abdullah Al Hafiz Khan, Sadid Hasan, Yong Pei
arXiv:2510. 16152v2 Announce Type: replace-cross Abstract: Scientific literature is increasingly fragmented by disciplinary boundaries, specialized terminology, and potentially sparse keyword systems, making it difficult to capture the evolving structure of modern science.
By Mason Smetana, Lev Khazanovich
BERTilda is an explainable framework for tracking topic lifecycles in longitudinal text streams. It discovers topics independently in each time window using an embedding‑based topic model, then links topics across adjacent windows via a temporal graph that uses both semantic similarity and a bidirectional coverage signal derived from tweet‑to‑topic attribution. The graph‑based rules identify continuations, splits, merges, disappearances, and unclear transitions, and the method achieves up to 87% agreement with human annotators on a gold‑standard subset.
By Cl\'audia Oliveira, \'Alvaro Figueira
arXiv:2603. 22510v2 Announce Type: replace-cross Abstract: Large language models are increasingly used in scholarly work, yet it remains unclear whether their productivity gains are accompanied by changes in research novelty.
By Ali Safari, Sahar Babaei
arXiv:2608. 13580v1 Announce Type: cross Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report.
By Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim, Mostafa Awad, Abdelrahman Sadallah, Gurpreet Gosal, Gokulakrishnan Ramakrishnan, Sarath Chandran, Biswajit Mishra, Rituraj Joshi, Ahmed Frikha, Etienne Goffinet, Abhishek Maiti, Ali El Filali, Sarah AlBarri, Samujjwal Ghosh, Rahul Pal, Parvez Mullah, Awantika Shukla, Sajid siddiki, Samta Kamboj, Onkar Pandit, Sunil Kumar Sahu, AbdelRahman Elbadawy, Amr Mohamed, Ahmad Chamma, Evan Dufraisse, Abdelaziz Bounhar, Dani Bouch, Hadi Abdine, Guokan Shang, Fajri Koto, Yuxia Wang, Zhuohan Xie, Ali Mekky, Rania Elbadry, Sarfraz Ahmad, Momina Ahsan, Omar El Herraoui, Daniil Orel, Hasan Iqbal, Kareem Elzeky, Mervat Abassy, Kareem Elozeiri, Saadeldine Eletter, Farah Atif, Nurdaulet Mukhituly, Haonan Li, Xudong Han, Aaryamonvikram Singh, Zainul Abedien Ahmed Quraishi, Neha Sengupta, Larry Murray, Avraham Sheinin, Joel Hestness, Natalia Vassilieva, Hector Xuguang Ren, Zhengzhong Liu, Michalis Vazirgiannis, Preslav Nakov
arXiv:2608. 19230v1 Announce Type: cross Abstract: As language models move from drafting prose to running literature-search agents with tool calls, fabricated references are becoming easier to catch and constrain.
By Sina Alemohammad, Denghui Zhang, Bolong Tang, Anthony Qin, Gengchen Mai, Ahmed Abbasi, Richard Baraniuk, Zhangyang Wang