arXiv Machine Learning By Benedikt Hilmes, Nick Rossenbach, Ralf Schl\"uter

Analyzing the Importance of Blank for CTC-Based Knowledge Distillation

Read the original on arXiv Machine Learning →

arXiv:2506. 01503v2 Announce Type: replace Abstract: With the rise of large pre-trained foundation models for automatic speech recognition new challenges appear.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jun 3

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance reconstruction quality with token efficiency, often encoding perceptually irrelevant information such as background noise and recording artifacts at the expense of linguistically and acoustically meaningful content.