arXiv:2606. 26144v1 Announce Type: cross Abstract: Speaker diarization, the task of determining "who spoke when" in a multi-speaker recording, is a critical component in applications such as meeting transcription, accessibility tools, and multilingual information retrieval.
By Samip Neupane, Sandesh Pokhrel, Sandesh Pyakurel, Basanta Joshi
arXiv:2607. 17164v1 Announce Type: new Abstract: Developing Automatic Speech Recognition (ASR) for morphologically rich, low-resource languages such as Assamese is challenging due to insufficient annotated speech data.
By Ganapati Das, Dwipen Laskar, Hasin Afzal Ahmed, Sanjib Kr Kalita, Kshirod Sarmah, Hem Chandra Das, Manjula Kalita
arXiv:2605. 13087v2 Announce Type: replace-cross Abstract: Fine-tuning multilingual ASR models like Whisper for low-resource languages often improves read speech but degrades spontaneous audio performance.
By Kush Juvekar, Kavya Manohar, Aditya Srinivas Menon, Arghya Bhattacharya, Kumarmanas Nethil
arXiv:2608.30092v1 Announce Type: cross
Abstract: We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single...
By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
arXiv:2609.14542v1 Announce Type: new
Abstract: Neyshekar is presented as an open Persian read-speech corpus designed for coverage of both formal and informal language, named entities, and longer utt...
By Ahmad Amirivojdan, Farzad Nadiri, Abolfazl Alizadeh, Shaghayegh Yaraghi
The paper investigates how to close the quality gap in low‑resource text‑to‑speech for Khmer and Korean using the VoxCPM2 model. By training a single low‑rank adaptation (LoRA) adapter on a shared 25.5‑hour corpus, the authors improve Khmer’s mean opinion score from 3.85 to 4.23 with a rank‑64 adapter, while Korean shows no significant gain. The study highlights that adaptation benefits mainly when the base model is weak and that training loss does not always align with human ratings.
By Phannet Pov, Hyun Woo Park, Voneat Pen, Sovandara Chhoun, Wan-Sup Cho, Saksonita Khoeurn
arXiv:2606. 24169v1 Announce Type: new Abstract: Adapting a streaming speech recognition model to a new language requires choosing between two plausible warm starts: a multilingual (ML) encoder or an English-only (EN) encoder.
By Nenad Banfic
arXiv:2609.14991v1 Announce Type: new
Abstract: Open Thai automatic speech recognition (ASR) is dominated by offline, Whisper-based models that read the whole utterance before transcribing, ruling ou...
By Warit Sirichotedumrong, Tanawin Samutsin, Shah Faisal Wani, Sittipong Sripaisarnmongkol, Kunat Pipatanakul
arXiv:2607. 23808v1 Announce Type: cross Abstract: In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India.
By Deovrat Mehendale, Aditya Mehndiratta, Dhruv Rathi, Kaushal Bhogale, Mitesh M. Khapra
SpeakPay is a voice‑first digital wallet designed to make mobile payment apps in Nepal accessible to visually impaired users. The paper introduces NepFinSpeech‑403, a 403‑utterance Nepali financial voice command dataset, and demonstrates that fine‑tuning Whisper large‑v2 with LoRA reduces the Word Error Rate from 129.95% to 42.58% and improves Devanagari numeral recognition from 0.0% to 73.9%. Domain adaptation also boosts the Transaction Success Rate from 1.67% to 33.33%, with as few as 100 domain‑specific utterances halving the zero‑shot WER.
By Biraj Subedi
The paper presents a method for creating a compact fixed‑voice Thai text‑to‑speech system by training a student model on synthetic speech generated from a large voice‑cloning teacher. By using only a short 15‑second voice reference and carefully filtering synthetic data, the authors build an 82‑million‑parameter model, Wayu‑Paxa‑TTS‑Edge, that runs on device without reference audio. The system achieves strong performance—68.2 % challenge‑set keyword accuracy, 91.4 % pause precision, and low character error rates—while outperforming its teacher and approaching the quality of a larger Gemini 3.1 model.
By Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut
BranchShine-CR is a 25‑million‑parameter model that transcribes multilingual speech into the International Phonetic Alphabet (IPA). It uses log‑mel features, a rotary‑position E‑Branchformer encoder, intermediate self‑conditioned CTC, and consistency regularization across augmented views. On 16,646 shared IPApack++ test utterances it achieves a 4.47 % IPA character error rate, a 22.3 % relative improvement over ZIPA‑CTC‑NS and outperforms a similarly sized NeMo Conformer baseline across all 41 language labels.
By Nikhil Navas, Sergio Chevtchenko, Talisson Damiao, Saeed Afshar