arXiv Computation and Language By V. S. D. S. Mahesh Akavarapu, Michael Daniel, Gerhard J\"ager

Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation

Read the original on arXiv Computation and Language →

The paper introduces two new corpus‑level, reference‑free metrics—Phoneme‑Cluster Mutual Information (PCMI) and Word Acoustic Consistency Score (WACS)—that use self‑supervised speech representations to evaluate forced alignment quality. PCMI quantifies how well aligned phoneme labels agree with clusters derived from SSL representations, while WACS assesses consistency across repeated word realizations via dynamic time warping of word representation sequences. Experiments on 85 languages from FLEURS and 45 languages in DoReCo show that both metrics degrade predictably under alignment perturbations, effectively distinguish high‑ from low‑quality alignments, and correlate strongly with traditional timestamp‑based measures, enabling scalable, multilingual alignment evaluation without manual annotations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.