Hugging Face Blog
Nov 15, 2021
The paper compares encoder‑based and generative decoder‑based large language models for evaluating automatic speech recognition (ASR). It examines BERTScore and SemDist across various LLMs, layers, and pooling strategies, finding that both metrics can strongly correlate with human judgments when properly configured. For generative LLMs, the study explores pairwise hypothesis selection via prompting and direct error classification, showing that while encoder‑based metrics remain competitive, generative models excel in hypothesis comparison and enhance interpretability of ASR evaluation.
arXiv:2607. 04011v1 Announce Type: cross Abstract: While decoders have rapidly scaled, encoders have remained largely unchanged since BERT.