arXiv AI By Akif Islam, Raufun Nahar, Md. Ekramul Hamid

When Audio Separation Hurts Zero-Shot ASR: Evaluating SAM-Audio with Whisper on Bengali and English Speech

Read the original on arXiv AI →

arXiv:2603. 04710v2 Announce Type: replace-cross Abstract: Recent advances in automatic speech recognition (ASR) and speech enhancement have strengthened the common belief that cleaner audio should lead to more accurate transcription.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 24

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

The paper introduces a synthetic Bengali speech dataset tailored for telecom customer‑care applications, comprising 10,000 audio‑text pairs (≈26.82 hours) with predefined train, validation, and test splits. The data were generated using OmniVoice voice‑cloning, and include both original and normalized transcripts for ASR/STT use. Automatic intelligibility evaluation with a fine‑tuned Whisper model shows an average WER of 2.54% and CER of 0.59%, indicating strong text‑audio consistency, while the authors note limitations of synthetic speech and STT‑based evaluation.

By Kawshik Kumar Paul, Md. Nafiul Alam Fuji
arXiv AI
Aug 20

Aslema at NADI 2026: Augmentation through Fewshot for SLU

Aslema is a system developed for the NADI 2026 Shared Task 5, which involves intent recognition and slot filling. The team evaluated four omni LLMs in a zero‑shot setting and found that fine‑tuned models consistently outperform zero‑shot inference. They further improved performance by augmenting data with culturally grounded Tunisian Derja utterances generated by an LLM and synthetic speech produced via voice cloning, achieving top‑ranked results on the official test set.

By Tajwaar Shafiq, Hunzalah Hassan Bhatti, Shammur Absar Chowdhury, Firoj Alam
arXiv Computation and Language
Sep 25

BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech

BanglaKontho is a newly released 20‑hour single‑speaker Bangla text‑to‑speech corpus derived from professional audiobook recordings, comprising 7,050 segmented utterances with verified transcripts at 24 kHz. The project also provides a reusable Bangla text normalizer that handles Bangladeshi‑style digit grouping, currency and date expressions, Danda punctuation, and Unicode normalization, along with the full preprocessing pipeline. A baseline MB‑iSTFT‑VITS model trained from scratch on this corpus achieves a 9.5 % WER and 4.46 naturalness MOS, outperforming the same architecture retrained on the smaller 12‑hour IndicTTS‑Bn corpus.

By Mizbaul Haque Maruf