EviSI is a large language model evaluation agent designed for simultaneous speech-to-speech translation. It adapts Multidimensional Quality Metrics to assess semantic fidelity and oral expression, using shared source evidence and deterministic scoring. In English‑to‑Chinese, EviSI achieves a mean Kendall agreement of 0.707 with human system rankings, outperforming baseline metrics, and shows positive concordance with COMET across five translation directions.
arXiv:2608. 08283v1 Announce Type: cross Abstract: Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation.
By Osvaldo Quinjica, Eric Bennett, Xinchen Yang, Andrew Schonebaum, Marine Carpuat
arXiv:2601.02933v4 Announce Type: replace
Abstract: Human evaluation is the gold standard for multilingual NLP, but is often skipped in practice and substituted with automatic metrics because it is n...
By Vil\'em Zouhar, Tom Kocmi
The paper introduces UPHELD, a large benchmark of human-to-human dialogues written by professional script writers, featuring realistic turn densities and over 36,000 per-turn human annotations. It evaluates existing automatic metrics and LLM-as-a-judge methods, finding them unreliable against expert human judgment. Using UPHELD, the authors develop a Mixture-of-Judges framework that improves correlation with human assessments by about 30%.
By Ilija Subasic, Andrew Rabinovich, Zhao Chen
arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.
By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran
Teochew has a substantial speaker community and exhibits distinctive lexical, syntactic, and pragmatic features, yet textual resources for evaluating large language models remain limited. We present T...