arXiv Computation and Language
Aug 27

Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

The paper compares encoder‑based and generative decoder‑based large language models for evaluating automatic speech recognition (ASR). It examines BERTScore and SemDist across various LLMs, layers, and pooling strategies, finding that both metrics can strongly correlate with human judgments when properly configured. For generative LLMs, the study explores pairwise hypothesis selection via prompting and direct error classification, showing that while encoder‑based metrics remain competitive, generative models excel in hypothesis comparison and enhance interpretability of ASR evaluation.

By Thibault Ba\~neras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso, Petr Motlicek, Shiran Liu, Mickael Rouvier, Jane Wottawa, Richard Dufour
arXiv Computation and Language
Sep 16

EviSI: An Evidence-Based Evaluation Agent for Simultaneous Interpreting

EviSI is an evidence‑based evaluation agent for low‑latency simultaneous speech‑to‑speech translation. It combines Multidimensional Quality Metrics with interpreter‑developed criteria, using shared source evidence to assess four dimensions—Anchor, Event, Logic, and Fluency—while deduplicating verified errors before scoring. On English‑to‑Chinese and Chinese‑to‑English data, EviSI’s rankings correlate strongly with human judgments, outperforming BLEU and COMET, and its multilingual extension maintains these correlations across five language directions.

By Ben Yan, Zongyao Li, Xiaoyu Chen, Daimeng Wei, Weidong Liu, Huan Zhao, Chong Li, Yaode Wang, Yuzhe Shang
Hugging Face Trending Papers
Sep 8

EviSI: An Evaluation Agent for Simultaneous Interpreting

EviSI is a large language model evaluation agent designed for simultaneous speech-to-speech translation. It adapts Multidimensional Quality Metrics to assess semantic fidelity and oral expression, using shared source evidence and deterministic scoring. In English‑to‑Chinese, EviSI achieves a mean Kendall agreement of 0.707 with human system rankings, outperforming baseline metrics, and shows positive concordance with COMET across five translation directions.

arXiv Machine Learning
Sep 21

Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR

The paper introduces a three‑stage pipeline to improve accented conversational ASR for speakers from India, Indonesia, and Latin America. It uses heuristic SQL filters to curate entity‑rich training data, regional LoRA adapters fine‑tuned on Qwen2.5‑Omni‑3B to generate both verbatim and corrected transcripts, and a six‑category error taxonomy validated by an LLM judge. The approach raises entity recall to 80‑85% and filler recall to 76‑86%, while keeping WER low (6‑10%) and outperforming Whisper and a commercial ASR on entity recall.

By Fiza Husain, Ankit Pandey, Yash Singh