arXiv Computation and Language

End-to-end Jordanian dialect speech-to-text self-supervised learning framework

arXiv Computation and Language
Sep 11

Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

Nuha‑Speech is a new initiative aimed at creating general‑purpose Arabic speech‑large language models (speech‑LLMs). It includes the construction of a large Arabic Speech Question‑Answering corpus with over 1.5 million samples for instruction tuning, supervised fine‑tuning of Qwen‑Omni model variants at various scales, and a systematic evaluation framework with diverse tasks and tailored metrics. The project seeks to establish foundational infrastructure for Arabic speech‑LLMs amid limited Arabic speech resources.

By Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi
arXiv AI
Jun 19

A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

arXiv:2606. 19747v1 Announce Type: new Abstract: Quran Automatic Speech Recognition (ASR) aims to convert Quranic recitation into text, enabling applications such as aided memorisation tools and Quranic search engines.

By Nabil Mosharraf Hossain (Greentech Apps Foundation, United Kingdom), Riasat Islam (Greentech Apps Foundation, United Kingdom, Queen Mary University of London, United Kingdom), Unaizah Obaidellah (University of Malaya, Malaysia)
arXiv AI
Aug 17

Jais 2: A Family of Arabic-Centric Open Large Language Models

arXiv:2608. 13580v1 Announce Type: cross Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report.

By Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim, Mostafa Awad, Abdelrahman Sadallah, Gurpreet Gosal, Gokulakrishnan Ramakrishnan, Sarath Chandran, Biswajit Mishra, Rituraj Joshi, Ahmed Frikha, Etienne Goffinet, Abhishek Maiti, Ali El Filali, Sarah AlBarri, Samujjwal Ghosh, Rahul Pal, Parvez Mullah, Awantika Shukla, Sajid siddiki, Samta Kamboj, Onkar Pandit, Sunil Kumar Sahu, AbdelRahman Elbadawy, Amr Mohamed, Ahmad Chamma, Evan Dufraisse, Abdelaziz Bounhar, Dani Bouch, Hadi Abdine, Guokan Shang, Fajri Koto, Yuxia Wang, Zhuohan Xie, Ali Mekky, Rania Elbadry, Sarfraz Ahmad, Momina Ahsan, Omar El Herraoui, Daniil Orel, Hasan Iqbal, Kareem Elzeky, Mervat Abassy, Kareem Elozeiri, Saadeldine Eletter, Farah Atif, Nurdaulet Mukhituly, Haonan Li, Xudong Han, Aaryamonvikram Singh, Zainul Abedien Ahmed Quraishi, Neha Sengupta, Larry Murray, Avraham Sheinin, Joel Hestness, Natalia Vassilieva, Hector Xuguang Ren, Zhengzhong Liu, Michalis Vazirgiannis, Preslav Nakov
arXiv Computation and Language
6d ago

Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models

The paper introduces a symbiotic architecture that equips large language models with audio‑understanding abilities without fine‑tuning their weights. It uses an injector module to write audio‑conditioned vectors into the LLM’s key‑value cache, allowing the model to act as an audio language model while keeping the backbone unchanged. The approach improves scalability—since injection cost depends on the injector width—and preserves the LLM’s original text performance, outperforming conventional frozen‑LLM methods and approaching fine‑tuned ALM results on audio tasks.

By Yotaro Kubo, Qi Sun, Yujin Tang
arXiv AI
Sep 21

MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs

MENASpeechBank is a new reference voice bank that provides about 18,000 high‑quality utterances from 124 speakers across multiple MENA countries, covering English, Modern Standard Arabic, and regional Arabic varieties. The dataset is built through a controllable pipeline that creates persona profiles inspired by the World Values Survey, defines a taxonomy of roughly 5,000 conversational scenarios, matches personas to scenarios via semantic similarity, and generates around 417,000 role‑play conversations using an LLM. Synthetic speaker‑conditioned user‑turn audio is produced from reference recordings to maintain speaker diversity, and both synthetic and human‑recorded conversations are evaluated and analyzed for quality.

By Zien Sheikh Ali, Hunzalah Hassan Bhatti, Rabindra Nath Nandi, Shammur Absar Chowdhury, Firoj Alam
arXiv Computation and Language
Sep 14

Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning

arXiv:2609.13045v1 Announce Type: new Abstract: Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-lingu...

By Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani, Alkis Koudounas, Rapha\"el Lafargue, Yosuke Kashiwagi, Quentin Jodelet, Emiru Tsunoo
Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.