The paper presents a method for creating a compact fixed‑voice Thai text‑to‑speech system by training a student model on synthetic speech generated from a large voice‑cloning teacher. By using only a short 15‑second voice reference and carefully filtering synthetic data, the authors build an 82‑million‑parameter model, Wayu‑Paxa‑TTS‑Edge, that runs on device without reference audio. The system achieves strong performance—68.2 % challenge‑set keyword accuracy, 91.4 % pause precision, and low character error rates—while outperforming its teacher and approaching the quality of a larger Gemini 3.1 model.
By Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut
arXiv:2607. 17164v1 Announce Type: new Abstract: Developing Automatic Speech Recognition (ASR) for morphologically rich, low-resource languages such as Assamese is challenging due to insufficient annotated speech data.
By Ganapati Das, Dwipen Laskar, Hasin Afzal Ahmed, Sanjib Kr Kalita, Kshirod Sarmah, Hem Chandra Das, Manjula Kalita
The Qwen-Audio-3.0-ASR Technical Report introduces a Mixture-of-Experts large language model-based automatic speech recognition system that addresses real‑world production challenges such as regional dialects, dynamic entities, hotwords, long‑range context, and disfluent speech. Built on the Qwen backbone and trained on tens of millions of hours of speech data, it supports transcription in 30 languages and 16 Chinese dialects, and offers industry‑domain entity recognition, hierarchical hotword customization, single‑pass polishing, and long‑audio contextual modeling. A streaming variant, Qwen-Audio-3.0-ASR-Streaming, is also presented for low‑latency applications, with evaluations showing state‑of‑the‑art performance against leading commercial systems.
By Chuanmeng Bian, Daren Chen, Peixin Chen, Zhigao Chen, Zhiyun Fan, Zhifu Gao, Bo Gong, Qing Gu, Jiajun He, Yawei Hu, Yunjie Ji, Jingbei Li, Xiangang Li, Xu Li, Zengxi Li, Zheng Li, Chengdong Liang, Baiji Liu, Ying Liu, Bin Ma, Yiping Peng, Yuezhang Peng, Zhendong Peng, Yu Pu, Yang Shi, Xin Shu, Jian Tang, Biao Tian, Peiyao Wang, Tianzi Wang, Wen Wang, Wupeng Wang, Cheng Wen, Yuzhong Wu, Zijian Xia, Yunchong Xiao, Nan Yang, Jianwei Yu, Jixing Yu, Binbin Zhang, Lei Zhang, Sitong Zhao, Guangdong Zhou, Yuan Zhou, Jianheng Zhuo
The paper introduces a new target‑speaker unlearning task for automatic speech recognition (TSU‑ASR) that allows certain speakers to opt out of transcription while still indicating their presence. A lightweight Enrollment‑Conditioned Gating (ECG) module is added to a frozen dual‑stream speech LLM, enabling dynamic unlearning of new opt‑out speakers during inference. Experiments on AMI and AliMeeting datasets show significant drops in transcription accuracy for opt‑out speakers while preserving performance for retained speakers.
By Bo Su, Yueru Yan, Thai Le
The paper introduces a synthetic Bengali speech dataset tailored for telecom customer‑care applications, comprising 10,000 audio‑text pairs (≈26.82 hours) with predefined train, validation, and test splits. The data were generated using OmniVoice voice‑cloning, and include both original and normalized transcripts for ASR/STT use. Automatic intelligibility evaluation with a fine‑tuned Whisper model shows an average WER of 2.54% and CER of 0.59%, indicating strong text‑audio consistency, while the authors note limitations of synthetic speech and STT‑based evaluation.
By Kawshik Kumar Paul, Md. Nafiul Alam Fuji
This study evaluates automatic speech recognition (ASR) for adolescent health communication in Twi, Dagbani, and Ewe by benchmarking five ASR systems on a Bible corpus and a domain-specific ASRH dataset, then performing supervised domain adaptation with a fine‑tuned Qwen3-ASR-0.6B model. Fine‑tuning significantly lowered word and character error rates, especially for Ewe, and the adapted model was deployed in the KasaHealth voice‑first application, which received high user approval and highlighted remaining domain gaps. The work demonstrates that in‑domain data, rather than model size or computational resources, is the primary limitation for effective ASR in these languages.
By Stephen E. Moore, Akwasi Asare, Mich-Seth Owusu, Paul Azunre, Joel Budu, Lawrence A. Adu-Gyamfi