arXiv AI

When Audio Separation Hurts Zero-Shot ASR: Evaluating SAM-Audio with Whisper on Bengali and English Speech

arXiv:2603. 04710v2 Announce Type: replace-cross Abstract: Recent advances in automatic speech recognition (ASR) and speech enhancement have strengthened the common belief that cleaner audio should lead to more accurate transcription.

arXiv Computation and Language
Aug 24

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

The paper introduces a synthetic Bengali speech dataset tailored for telecom customer‑care applications, comprising 10,000 audio‑text pairs (≈26.82 hours) with predefined train, validation, and test splits. The data were generated using OmniVoice voice‑cloning, and include both original and normalized transcripts for ASR/STT use. Automatic intelligibility evaluation with a fine‑tuned Whisper model shows an average WER of 2.54% and CER of 0.59%, indicating strong text‑audio consistency, while the authors note limitations of synthetic speech and STT‑based evaluation.

By Kawshik Kumar Paul, Md. Nafiul Alam Fuji
arXiv AI
Aug 20

Aslema at NADI 2026: Augmentation through Fewshot for SLU

Aslema is a system developed for the NADI 2026 Shared Task 5, which involves intent recognition and slot filling. The team evaluated four omni LLMs in a zero‑shot setting and found that fine‑tuned models consistently outperform zero‑shot inference. They further improved performance by augmenting data with culturally grounded Tunisian Derja utterances generated by an LLM and synthetic speech produced via voice cloning, achieving top‑ranked results on the official test set.

By Tajwaar Shafiq, Hunzalah Hassan Bhatti, Shammur Absar Chowdhury, Firoj Alam
arXiv Computation and Language
Sep 25

BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech

BanglaKontho is a newly released 20‑hour single‑speaker Bangla text‑to‑speech corpus derived from professional audiobook recordings, comprising 7,050 segmented utterances with verified transcripts at 24 kHz. The project also provides a reusable Bangla text normalizer that handles Bangladeshi‑style digit grouping, currency and date expressions, Danda punctuation, and Unicode normalization, along with the full preprocessing pipeline. A baseline MB‑iSTFT‑VITS model trained from scratch on this corpus achieves a 9.5 % WER and 4.46 naturalness MOS, outperforming the same architecture retrained on the smaller 12‑hour IndicTTS‑Bn corpus.

By Mizbaul Haque Maruf
arXiv AI
Aug 25

Aslema at NADI 2026: Data Augmentation for Intent Recognition and Slot Filling

The paper introduces Aslema, a system for the NADI 2026 Shared Task 5, which includes intent recognition and slot filling. The authors evaluate four omni LLMs in zero‑shot and fine‑tuned settings, finding that fine‑tuning consistently outperforms zero‑shot inference. They further augment data by generating culturally grounded Tunisian Derja utterances with an LLM and synthetic speech via voice cloning, which improves performance; the final system based on Qwen3‑Omni‑30B achieves 86.8% intent accuracy and 34.7 WER on devtest, ranking 1st in slot filling and 4th in intent recognition on the official test set.

By Tajwaar Shafiq, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
arXiv AI
Sep 18

Model-Agnostic and Language-Agnostic Voice Pipeline Improvement for the Agriculture Domain

The paper introduces a modular, model‑agnostic pipeline to enhance automatic speech recognition for FarmerChat, an AI agricultural advisory assistant used by smallholder farmers in their native languages. The pipeline integrates gated audio enhancement, speaker diarization with target‑speaker selection, domain‑aware lexicon correction, and a quality gate, requiring fine‑tuning only at the diarization stage. Evaluations on Hindi, Telugu, and Odia recordings show significant reductions in word error rate—up to 42% on multi‑speaker cloud ASR models—demonstrating that targeted preprocessing and domain‑specific post‑processing can markedly improve transcription quality without altering the core ASR model.

By Aakash Singh, Lakshmi Pedapudi, Chandrashekar M S, Sanyam Singh, Naga Ganesh, Vineet Singh
arXiv Computation and Language
Sep 25

BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech

BanglaTurn is a new corpus of 35,374 Bangla podcast speech samples, each 3 to 15 seconds long, labeled for end‑of‑turn detection through speaker diarization, an LLM pass, and human verification. A Whisper‑based model with task‑specific classification heads achieves 84.33 % accuracy on a balanced test set, outperforming the Smart‑Turn v3 baseline (69.28 %) and reducing the false‑negative rate from 51.57 % to 7.55 %, though with a higher false‑positive rate. The study also details the contributions of encoder‑layer fine‑tuning, multi‑scale pooling, INT8 quantization, and reports inference latency of 165–191 ms on CPU.

By Mizbaul Haque Maruf