arXiv Computation and Language

T-SANDHI: Tone Sandhi-aware Adaptive Network with Decoupled Hybrid Injection for Low-resource Taiwanese Hokkien Speech Recognition

T‑SANDHI is a lightweight Taiwanese Hokkien ASR model that addresses tone sandhi by decoupling surface acoustics from lexical intent on a frozen Whisper backbone. It uses a lexicon‑guided multi‑task learning framework with text‑derived pseudo labels and a hybrid injection module that dynamically gates independent citation and sandhi phonetic streams. Experiments on the TAT‑MOE corpus and two blind test sets show that this explicit disentanglement resolves tonal mapping confusion and outperforms baselines while keeping parameter usage low.

arXiv AI
Sep 4

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

The paper presents a method for creating a compact fixed‑voice Thai text‑to‑speech system by training a student model on synthetic speech generated from a large voice‑cloning teacher. By using only a short 15‑second voice reference and carefully filtering synthetic data, the authors build an 82‑million‑parameter model, Wayu‑Paxa‑TTS‑Edge, that runs on device without reference audio. The system achieves strong performance—68.2 % challenge‑set keyword accuracy, 91.4 % pause precision, and low character error rates—while outperforming its teacher and approaching the quality of a larger Gemini 3.1 model.

By Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut
Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.

arXiv Computation and Language
Sep 25

VietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching

VietPrism is a newly released, large‑scale Vietnamese speech corpus that combines 993.4 hours of real utterances from 1,262 verified speakers with 3.1 k hours of synthetic spoof speech. It uniquely offers transcripts, consistent speaker identities, five dialect groups, and extensive Vietnamese‑English code‑switching—nearly half of the corpus—while pairing each spoof with a matched bona fide utterance. The dataset enables controlled evaluation of deep‑fake detection models, revealing significant variability in detector performance across dialects and speaker similarity.

By Minh Hoang, Thai Le