arXiv Computation and Language

Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge

The paper introduces Muslim, an Arabic voice AI platform that delivers grounded Islamic knowledge to users in real time. It combines a NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, and a self-hosted TTS system, supported by a deterministic multi-source retrieval layer across six Model Context Protocol servers. The authors release fine‑tuned Arabic Islamic model artifacts, describe an account‑based metering layer that prevents abuse, and present a three‑layer observability stack that monitors GPU‑bound agent hosts, reporting latency, accuracy, and engineering trade‑offs for production deployment.

arXiv AI
Aug 24

Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations

Ansari is a retrieval‑grounded Islamic AI assistant that has handled over 140,000 conversations in more than 25 languages since June 2023. It uses an agentic retrieval loop where a language model searches authenticated Islamic corpora—including the Qur’an, hadith collections, fiqh encyclopedias, and tafsir sources—and answers only based on retrieved content, providing citations for verification. The paper details Ansari’s architecture, multi‑platform deployment, evaluation results (including top performance on the IslamicMMLU leaderboard and strong resistance to false premises), and lessons for faith‑sensitive LLM deployments.

By M Waleed Kadous, Amr Elsayed, Abdullah Al Nahas, Ashraf Haress
arXiv AI
Aug 17

Jais 2: A Family of Arabic-Centric Open Large Language Models

arXiv:2608. 13580v1 Announce Type: cross Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report.

By Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim, Mostafa Awad, Abdelrahman Sadallah, Gurpreet Gosal, Gokulakrishnan Ramakrishnan, Sarath Chandran, Biswajit Mishra, Rituraj Joshi, Ahmed Frikha, Etienne Goffinet, Abhishek Maiti, Ali El Filali, Sarah AlBarri, Samujjwal Ghosh, Rahul Pal, Parvez Mullah, Awantika Shukla, Sajid siddiki, Samta Kamboj, Onkar Pandit, Sunil Kumar Sahu, AbdelRahman Elbadawy, Amr Mohamed, Ahmad Chamma, Evan Dufraisse, Abdelaziz Bounhar, Dani Bouch, Hadi Abdine, Guokan Shang, Fajri Koto, Yuxia Wang, Zhuohan Xie, Ali Mekky, Rania Elbadry, Sarfraz Ahmad, Momina Ahsan, Omar El Herraoui, Daniil Orel, Hasan Iqbal, Kareem Elzeky, Mervat Abassy, Kareem Elozeiri, Saadeldine Eletter, Farah Atif, Nurdaulet Mukhituly, Haonan Li, Xudong Han, Aaryamonvikram Singh, Zainul Abedien Ahmed Quraishi, Neha Sengupta, Larry Murray, Avraham Sheinin, Joel Hestness, Natalia Vassilieva, Hector Xuguang Ren, Zhengzhong Liu, Michalis Vazirgiannis, Preslav Nakov
arXiv AI
Sep 7

Scalable Context Orchestration for Serving LLMs Over Voice

The paper introduces llmovoice, a middleware that explicitly models voice context for large language model (LLM) serving in voice AI applications. By incorporating speaking rate, background noise, packet loss, and other paralinguistic factors into a bounded context, llmovoice guides the LLM to generate more aligned responses. Experiments show significant reductions in speaking‑rate errors, false interruptions, and model usage costs, especially in long voice sessions.

By Linyi Jiang, Silvery D. Fu, Yifei Zhu
Hugging Face Trending Papers
Sep 3

Scalable Context Orchestration for Serving LLMs Over Voice

Scalable Context Orchestration for Serving LLMs Over Voice presents llmovoice, a middleware that explicitly models voice context—including speaking rate, background noise, and packet loss—to guide large language model responses. By constructing a bounded voice context at each turn, llmovoice improves alignment with user preferences and reduces errors, achieving a 52.4% drop in speaking‑rate alignment error and a 0.9% false‑interruption rate under packet loss. In addition, it cuts model usage costs dramatically, lowering per‑turn cost by up to 24.9× while maintaining 98.7% of baseline answer quality in long sessions.

arXiv Computation and Language
5d ago

TutlAit v1: a crowdsourced Moroccan Tamazight speech dataset with Arabic transcriptions and regional accent labels

The TutlAit v1 dataset is a crowdsourced corpus of Moroccan Tamazight speech paired with Modern Standard Arabic transcriptions and explicit regional accent labels. It contains 13,384 audio files (≈20.9 hours) collected via a web application, with volunteers contributing through text‑to‑audio and audio‑to‑text workflows, and includes additional segments from freely available media. The dataset covers Atlas, Souss, Rif, and Kabyle varieties, making it a valuable resource for speech recognition, translation, and accent identification in an under‑resourced language.

By Mohamed-Amine Chadi, Ezzahra Ait El Arbi, Ismail Khayoub, Aymane Fadili, Yassine Ennhili, Jadjigua Bouali, Hanane Inhid, Mohammed Ameksa, Hajar Mousannif
arXiv Computation and Language
Sep 3

SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition

SpeakPay is a voice‑first digital wallet designed to make mobile payment apps in Nepal accessible to visually impaired users. The paper introduces NepFinSpeech‑403, a 403‑utterance Nepali financial voice command dataset, and demonstrates that fine‑tuning Whisper large‑v2 with LoRA reduces the Word Error Rate from 129.95% to 42.58% and improves Devanagari numeral recognition from 0.0% to 73.9%. Domain adaptation also boosts the Transaction Success Rate from 1.67% to 33.33%, with as few as 100 domain‑specific utterances halving the zero‑shot WER.

By Biraj Subedi
arXiv AI
Sep 21

MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs

MENASpeechBank is a new reference voice bank that provides about 18,000 high‑quality utterances from 124 speakers across multiple MENA countries, covering English, Modern Standard Arabic, and regional Arabic varieties. The dataset is built through a controllable pipeline that creates persona profiles inspired by the World Values Survey, defines a taxonomy of roughly 5,000 conversational scenarios, matches personas to scenarios via semantic similarity, and generates around 417,000 role‑play conversations using an LLM. Synthetic speaker‑conditioned user‑turn audio is produced from reference recordings to maintain speaker diversity, and both synthetic and human‑recorded conversations are evaluated and analyzed for quality.

By Zien Sheikh Ali, Hunzalah Hassan Bhatti, Rabindra Nath Nandi, Shammur Absar Chowdhury, Firoj Alam
arXiv Machine Learning
Sep 29

LUMO (Lightweight Unified Multilingual Orchestrator): A Privacy Preserving Offline Voice Assistant

LUMO (Lightweight Unified Multilingual Orchestrator) is a privacy‑preserving offline voice assistant that runs entirely on edge hardware, specifically a Raspberry Pi 5 with 8 GB RAM. It integrates local ASR, a 4‑bit GGUF‑quantized LLM, and TTS to deliver end‑to‑end response latencies of 2.0–4.0 s, a 6.8 % WER on short English utterances, and lower peak power consumption (~9 W) compared to existing edge assistants. The system also supports Bangla speech, enabling multilingual use in low‑resource settings.

By Md. Mehedi Hasan Naeem, Mst. Kamrunnahar Ruma, Nafiza Anjum, Shakila Sultana, Md. Sujan Ali
arXiv AI
Aug 12

MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model

arXiv:2608. 10360v1 Announce Type: cross Abstract: Arabic maqam music microtonal, modal, and built on ornamented call and response is among the traditions most underserved by generative music models, whose training frameworks remain predominantly Western and equaltempered.

By Jiaxin Du, Boulbaba Abdeljaouad, Yong Zhuang, Haoyu Li