arXiv Computation and Language By Lawry Sorenson, Michael Crandall, Eric K. Ringger, Stephen D. Richardson

Scaling Forced Alignment to End-User Devices

Read the original on arXiv Computation and Language →

The paper presents two optimizations for forced alignment of audio to text using the Viterbi algorithm. The first optimization applies the Hirschberg algorithm to reduce memory usage from 140 GB to 5 MB for three‑hour inputs and speeds up alignment to one‑third the time of torchaudio on a CPU. The second optimization models alignment as a constrained random walk, enabling pruning that yields an additional 2× speedup on inputs longer than 20 minutes while maintaining over 98 % alignment accuracy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.