Hugging Face Trending Papers

Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting

Read the original on Hugging Face Trending Papers →

The paper introduces a new corpus of 1,101 human simultaneous interpreting segments, each scored on meaning transfer, delivery quality, and perceived latency. It demonstrates that standard large‑language‑model prompting and scalar supervision fail to align with these rubric dimensions, producing near‑zero correlation with human ratings. By adding dual regression heads to a LoRA‑adapted COMET‑KIWI encoder, the authors achieve modest Pearson correlations (0.388 for meaning transfer and 0.301 for delivery quality) on a held‑out test set, improving over the frozen baseline.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
Sep 11

Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting

The paper introduces a new automatic evaluation metric for human simultaneous interpreting that aligns with traditional analytic rubrics. A corpus of 1,101 professionally scored SI segments is created, covering meaning transfer, delivery quality, and perceived latency. Using a LoRA‑adapted COMET‑KIWI encoder with dual regression heads, the model achieves Pearson correlations of 0.388 for meaning transfer and 0.301 for delivery quality, outperforming the frozen baseline while acknowledging low rater agreement.

By Ziyu Zhang, Satoshi Nakamura
arXiv Computation and Language
Sep 16

EviSI: An Evidence-Based Evaluation Agent for Simultaneous Interpreting

EviSI is an evidence‑based evaluation agent for low‑latency simultaneous speech‑to‑speech translation. It combines Multidimensional Quality Metrics with interpreter‑developed criteria, using shared source evidence to assess four dimensions—Anchor, Event, Logic, and Fluency—while deduplicating verified errors before scoring. On English‑to‑Chinese and Chinese‑to‑English data, EviSI’s rankings correlate strongly with human judgments, outperforming BLEU and COMET, and its multilingual extension maintains these correlations across five language directions.

By Ben Yan, Zongyao Li, Xiaoyu Chen, Daimeng Wei, Weidong Liu, Huan Zhao, Chong Li, Yaode Wang, Yuzhe Shang
Hugging Face Trending Papers
Sep 8

EviSI: An Evaluation Agent for Simultaneous Interpreting

EviSI is a large language model evaluation agent designed for simultaneous speech-to-speech translation. It adapts Multidimensional Quality Metrics to assess semantic fidelity and oral expression, using shared source evidence and deterministic scoring. In English‑to‑Chinese, EviSI achieves a mean Kendall agreement of 0.707 with human system rankings, outperforming baseline metrics, and shows positive concordance with COMET across five translation directions.