arXiv:2607. 23808v1 Announce Type: cross Abstract: In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India.
By Deovrat Mehendale, Aditya Mehndiratta, Dhruv Rathi, Kaushal Bhogale, Mitesh M. Khapra
The paper introduces a synthetic Bengali speech dataset tailored for telecom customer‑care applications, comprising 10,000 audio‑text pairs (≈26.82 hours) with predefined train, validation, and test splits. The data were generated using OmniVoice voice‑cloning, and include both original and normalized transcripts for ASR/STT use. Automatic intelligibility evaluation with a fine‑tuned Whisper model shows an average WER of 2.54% and CER of 0.59%, indicating strong text‑audio consistency, while the authors note limitations of synthetic speech and STT‑based evaluation.
By Kawshik Kumar Paul, Md. Nafiul Alam Fuji
arXiv:2606. 26901v1 Announce Type: cross Abstract: Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown.
By Subham Kumar, Prakrithi Shivaprakash, Abhishek Manoharan, Astut Kurariya, Diptadhi Mukherjee, Prabhat Chand, Pratima Murthy, Koustav Rudra, Lekhansh Shukla, Animesh Mukherjee
arXiv:2605. 13087v2 Announce Type: replace-cross Abstract: Fine-tuning multilingual ASR models like Whisper for low-resource languages often improves read speech but degrades spontaneous audio performance.
By Kush Juvekar, Kavya Manohar, Aditya Srinivas Menon, Arghya Bhattacharya, Kumarmanas Nethil
arXiv:2608. 12327v1 Announce Type: cross Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol.
By Suman Paudel, Sarbin Sayami
The paper evaluates bias in phoneme-based automatic speech recognition (ASR) systems, focusing on WhisperIPA and ZIPA, which produce International Phonetic Alphabet (IPA) transcriptions. Using multilingual speech corpora and demographically annotated English datasets, the authors compare model-generated IPA against grapheme-to-phoneme (G2P) outputs with both standard phoneme error rate (PER) and a new Soft PER metric that allows linguistically similar substitutions. The study finds persistent disparities across language, gender, accent, ethnicity, and age, even when accounting for acceptable phonemic variation.
By Maneesha Rani Saha, Catherine Bao, Neal Patwari
arXiv:2606. 03504v1 Announce Type: cross Abstract: We present BaltiVoice, a 16.
By Muhammad Ali
arXiv:2605. 20712v2 Announce Type: replace-cross Abstract: Automatic speech recognition replaces typing only when correction costs less than manual entry - a threshold determined by error types, not counts: fixing a misrecognized domain term costs far more than inserting a comma.
By Kavya Manohar, Arghya Bhattacharya, Kush Juvekar, Kumarmanas Nethil
The study fine‑tunes the Whisper Small model for Automatic Speech Recognition (ASR) in Baniwa, an indigenous Arawakan language. Using a 0.54‑hour corpus of 1,373 manually transcribed recordings, the fine‑tuned model achieved a Word Error Rate of 37.5% and a Character Error Rate of 7.45%. These results provide an initial baseline for Baniwa ASR and suggest that multilingual foundation models can be adapted to extremely low‑resource languages.
By Leonardo Duart, Tiago Fonseca, Thiago Chac\'on
arXiv:2603. 04710v2 Announce Type: replace-cross Abstract: Recent advances in automatic speech recognition (ASR) and speech enhancement have strengthened the common belief that cleaner audio should lead to more accurate transcription.
By Akif Islam, Raufun Nahar, Md. Ekramul Hamid
arXiv:2606. 17826v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) in non-English clinical settings is challenged by multiscript variability, where the same term may appear in multiple valid orthographic forms.
By Jean Seo, Minkyu Kim, Jeonguk Lee, Jisoo Jung, Wooseok Han, Eunho Yang
We present BaltiVoice, a 16. 8-hour read-speech corpus for Balti (ISO 639-3: bft), a Tibetic language spoken in Gilgit-Baltistan, Pakistan, with no prior publicly available ASR resources.