arXiv AI By Siddhant Hitesh Mantri, Dhara Gorasiya, Malhar Kulkarni, Pushpak Bhattacharya

From Lexicon to AI: A Structured-Data Pipeline for Specialized Conversational Systems in Low-Resource Languages

Read the original on arXiv AI →

arXiv:2606. 26112v1 Announce Type: cross Abstract: Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive training corpora.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations

The paper introduces Human-1, the first open and reproducible full‑duplex spoken dialogue system for Hindi. It adapts the Moshi architecture with a custom Hindi tokenizer and trains on 26,000 hours of real spontaneous conversations from 14,695 speakers, enabling the model to learn natural turn‑taking, interruptions, overlaps, and backchannels. A two‑stage training process—large‑scale pre‑training followed by fine‑tuning on 1,000 hours—yields a system that, according to both automatic metrics and human judgments, generates natural and meaningful full‑duplex conversational behavior in Hindi.

By Bhaskar Singh, Manas Dhir, Shobhit Banga, Pranav Sharma
arXiv Computation and Language
Aug 28

DocTalkBN: A Novel Dataset of Expert Telemedicine Conversations in Bengali

DocTalkBN is a large-scale multimodal dataset of authentic expert telemedicine conversations in Bengali, comprising 557.63 hours of paired audio and text, 1,515 multi-turn patient calls, and 10,274 host–doctor question–answer exchanges across 26 medical specialties. The dataset contains 1.7 million tokens and preserves the spontaneity and contextual richness of real medical interactions in a low-resource language. Three downstream tasks—medical triage classification, advice safety evaluation, and medical named entity recognition—are constructed to benchmark large language models and encoder-based baselines, demonstrating DocTalkBN’s practical usefulness for clinically grounded reasoning.

By Anik Saha, Fahmida Sultana Naznin, Sadatul Islam Sadi, Ananya Shahrin Promi, Wahid Al Azad Navid, Rifat Shahriyar
arXiv Machine Learning
Jul 28

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

arXiv:2607. 23242v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms.

By Sahil Deepak Gawande, Mayank Singh
arXiv Computation and Language
Sep 11

Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

Nuha‑Speech is a new initiative aimed at creating general‑purpose Arabic speech‑large language models (speech‑LLMs). It includes the construction of a large Arabic Speech Question‑Answering corpus with over 1.5 million samples for instruction tuning, supervised fine‑tuning of Qwen‑Omni model variants at various scales, and a systematic evaluation framework with diverse tasks and tailored metrics. The project seeks to establish foundational infrastructure for Arabic speech‑LLMs amid limited Arabic speech resources.

By Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi