arXiv Computation and Language
Sep 17

T-SANDHI: Tone Sandhi-aware Adaptive Network with Decoupled Hybrid Injection for Low-resource Taiwanese Hokkien Speech Recognition

T‑SANDHI is a lightweight Taiwanese Hokkien ASR model that addresses tone sandhi by decoupling surface acoustics from lexical intent on a frozen Whisper backbone. It uses a lexicon‑guided multi‑task learning framework with text‑derived pseudo labels and a hybrid injection module that dynamically gates independent citation and sandhi phonetic streams. Experiments on the TAT‑MOE corpus and two blind test sets show that this explicit disentanglement resolves tonal mapping confusion and outperforms baselines while keeping parameter usage low.

By Hung-Yang Sung, Chien-Chun Wang, Tien-Hong Lo, Yu-Sheng Tsao, Yung-Chang Hsu, Berlin Chen