arXiv Computation and Language By Mizbaul Haque Maruf

BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech

Read the original on arXiv Computation and Language →

BanglaKontho is a newly released 20‑hour single‑speaker Bangla text‑to‑speech corpus derived from professional audiobook recordings, comprising 7,050 segmented utterances with verified transcripts at 24 kHz. The project also provides a reusable Bangla text normalizer that handles Bangladeshi‑style digit grouping, currency and date expressions, Danda punctuation, and Unicode normalization, along with the full preprocessing pipeline. A baseline MB‑iSTFT‑VITS model trained from scratch on this corpus achieves a 9.5 % WER and 4.46 naturalness MOS, outperforming the same architecture retrained on the smaller 12‑hour IndicTTS‑Bn corpus.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Aug 24

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

The paper introduces a synthetic Bengali speech dataset tailored for telecom customer‑care applications, comprising 10,000 audio‑text pairs (≈26.82 hours) with predefined train, validation, and test splits. The data were generated using OmniVoice voice‑cloning, and include both original and normalized transcripts for ASR/STT use. Automatic intelligibility evaluation with a fine‑tuned Whisper model shows an average WER of 2.54% and CER of 0.59%, indicating strong text‑audio consistency, while the authors note limitations of synthetic speech and STT‑based evaluation.

By Kawshik Kumar Paul, Md. Nafiul Alam Fuji
arXiv Computation and Language
Sep 25

BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech

BanglaTurn is a new corpus of 35,374 Bangla podcast speech samples, each 3 to 15 seconds long, labeled for end‑of‑turn detection through speaker diarization, an LLM pass, and human verification. A Whisper‑based model with task‑specific classification heads achieves 84.33 % accuracy on a balanced test set, outperforming the Smart‑Turn v3 baseline (69.28 %) and reducing the false‑negative rate from 51.57 % to 7.55 %, though with a higher false‑positive rate. The study also details the contributions of encoder‑layer fine‑tuning, multi‑scale pooling, INT8 quantization, and reports inference latency of 165–191 ms on CPU.

By Mizbaul Haque Maruf