arXiv:2608.28611v1 Announce Type: cross
Abstract: Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly trai...
By Isha Narang, Sneh Gosai, Mayank Singh
The paper introduces Human-1, the first open and reproducible full‑duplex spoken dialogue system for Hindi. It adapts the Moshi architecture with a custom Hindi tokenizer and trains on 26,000 hours of real spontaneous conversations from 14,695 speakers, enabling the model to learn natural turn‑taking, interruptions, overlaps, and backchannels. A two‑stage training process—large‑scale pre‑training followed by fine‑tuning on 1,000 hours—yields a system that, according to both automatic metrics and human judgments, generates natural and meaningful full‑duplex conversational behavior in Hindi.
By Bhaskar Singh, Manas Dhir, Shobhit Banga, Pranav Sharma
DocTalkBN is a large-scale multimodal dataset of authentic expert telemedicine conversations in Bengali, comprising 557.63 hours of paired audio and text, 1,515 multi-turn patient calls, and 10,274 host–doctor question–answer exchanges across 26 medical specialties. The dataset contains 1.7 million tokens and preserves the spontaneity and contextual richness of real medical interactions in a low-resource language. Three downstream tasks—medical triage classification, advice safety evaluation, and medical named entity recognition—are constructed to benchmark large language models and encoder-based baselines, demonstrating DocTalkBN’s practical usefulness for clinically grounded reasoning.
By Anik Saha, Fahmida Sultana Naznin, Sadatul Islam Sadi, Ananya Shahrin Promi, Wahid Al Azad Navid, Rifat Shahriyar
arXiv:2607. 23242v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms.
By Sahil Deepak Gawande, Mayank Singh
Nuha‑Speech is a new initiative aimed at creating general‑purpose Arabic speech‑large language models (speech‑LLMs). It includes the construction of a large Arabic Speech Question‑Answering corpus with over 1.5 million samples for instruction tuning, supervised fine‑tuning of Qwen‑Omni model variants at various scales, and a systematic evaluation framework with diverse tasks and tailored metrics. The project seeks to establish foundational infrastructure for Arabic speech‑LLMs amid limited Arabic speech resources.
By Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi
arXiv:2609.24410v1 Announce Type: new
Abstract: Speech-to-text engines are extremely needed nowadays for different applications, representing an essential enabler in human-robot interaction. Still, s...
By Ali A. Safieh, Ibrahim Abu Alhaol, Rawan Ghnemat
arXiv:2608. 13580v1 Announce Type: cross Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report.
By Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim, Mostafa Awad, Abdelrahman Sadallah, Gurpreet Gosal, Gokulakrishnan Ramakrishnan, Sarath Chandran, Biswajit Mishra, Rituraj Joshi, Ahmed Frikha, Etienne Goffinet, Abhishek Maiti, Ali El Filali, Sarah AlBarri, Samujjwal Ghosh, Rahul Pal, Parvez Mullah, Awantika Shukla, Sajid siddiki, Samta Kamboj, Onkar Pandit, Sunil Kumar Sahu, AbdelRahman Elbadawy, Amr Mohamed, Ahmad Chamma, Evan Dufraisse, Abdelaziz Bounhar, Dani Bouch, Hadi Abdine, Guokan Shang, Fajri Koto, Yuxia Wang, Zhuohan Xie, Ali Mekky, Rania Elbadry, Sarfraz Ahmad, Momina Ahsan, Omar El Herraoui, Daniil Orel, Hasan Iqbal, Kareem Elzeky, Mervat Abassy, Kareem Elozeiri, Saadeldine Eletter, Farah Atif, Nurdaulet Mukhituly, Haonan Li, Xudong Han, Aaryamonvikram Singh, Zainul Abedien Ahmed Quraishi, Neha Sengupta, Larry Murray, Avraham Sheinin, Joel Hestness, Natalia Vassilieva, Hector Xuguang Ren, Zhengzhong Liu, Michalis Vazirgiannis, Preslav Nakov
arXiv:2608. 04586v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT).
By Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Chengpeng Fu, Yu Wang, Ming Liu
arXiv:2606. 03957v1 Announce Type: cross Abstract: Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data.
By M\'at\'e Gedeon, P\'eter Mihajlik
arXiv:2608. 04586v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT).
By Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Chengpeng Fu, Yu Wang, Ming Liu
The paper introduces an LLM-based Conversational AI Knowledge Assistant for the Raspberry‑Pi‑powered 13‑Axis MyBuddy humanoid robot. It combines large language model-driven language understanding, real‑time speech recognition, internet‑based knowledge retrieval (e.g., Wikipedia, arXiv), flexible dialogue management, and natural speech synthesis to support intelligent, multi‑turn conversations and emotional‑support interactions. This system aims to overcome the limitations of traditional rule‑based dialogue systems in humanoid robots.
By Hanxiao Chen
arXiv:2606. 24825v1 Announce Type: cross Abstract: Part-of-Speech (POS) tagging is a foundational NLP task underpinning machine translation, information extraction, and syntactic parsing.
By Hariom Ingle, Ronit Ghode, Ishwari Gondkar, Jidnyasa Harad, Raviraj Joshi