arXiv Computation and Language

Scaling Forced Alignment to End-User Devices

The paper presents two optimizations for forced alignment of audio to text using the Viterbi algorithm. The first optimization applies the Hirschberg algorithm to reduce memory usage from 140 GB to 5 MB for three‑hour inputs and speeds up alignment to one‑third the time of torchaudio on a CPU. The second optimization models alignment as a constrained random walk, enabling pruning that yields an additional 2× speedup on inputs longer than 20 minutes while maintaining over 98 % alignment accuracy.

arXiv Machine Learning
Jul 20

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching

arXiv:2605. 22083v2 Announce Type: replace-cross Abstract: While flow-matching text-to-speech (TTS) achieves strong zero-shot speaker similarity and naturalness, it remains susceptible to content fidelity issues, particularly skip and repeat errors from imperfect alignment.

By Jinhyeok Yang, Hyeongju Kim, Yechan Yu, Joon Byun, Frederik Bous, Juheon Lee