AskQE: Question Answering as Automatic Evaluation for Machine Translation
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
EviSI is an evidence‑based evaluation agent for low‑latency simultaneous speech‑to‑speech translation. It combines Multidimensional Quality Metrics with interpreter‑developed criteria, using shared source evidence to assess four dimensions—Anchor, Event, Logic, and Fluency—while deduplicating verified errors before scoring. On English‑to‑Chinese and Chinese‑to‑English data, EviSI’s rankings correlate strongly with human judgments, outperforming BLEU and COMET, and its multilingual extension maintains these correlations across five language directions.
The paper introduces Iterative MBR Distillation for Error Span Detection (ESD) in machine translation, a self‑evolution framework that replaces human annotations with pseudo‑labels generated by a large language model. By iteratively applying Minimum Bayes Risk decoding, the method produces high‑quality error spans without costly human effort. Experiments on WMT Metrics Shared Task datasets show that models trained solely on these pseudo‑labels outperform both unadapted baselines and supervised models trained on human data at system and span levels, while keeping sentence‑level performance competitive.
The paper introduces the Last Translation Benchmark (LTB), a live dataset of human-authored and peer‑reviewed examples—including texts, images, audio, and videos—that are designed to break current state‑of‑the‑art machine translation models. Each example is accompanied by handcrafted verification rules that specify concrete failure cases, providing a reliable and actionable evaluation method. The benchmark aims to overcome the limitations of existing automatic metrics and gold human evaluations, which often lack reproducibility, objectivity, and scalability.
arXiv:2609.22793v1 Announce Type: new Abstract: LLM-based machine translation evaluation can closely match human judgments, but in practice it remains largely diagnostic, with the signals rarely tran...
EDRAC is the first large‑scale benchmark for dialectal Arabic machine reading comprehension and generative question answering, covering five major dialects—Egyptian, Moroccan, Emirati, Syrian, and Saudi. It contains 499 passages from naturally spoken interactions and 4,977 QA pairs produced via a human–LLM collaborative pipeline. The benchmark evaluates Arabic‑centric and multilingual large language models, revealing gaps between semantic answer quality and dialectal fidelity and underscoring limitations of current evaluation metrics for dialectal Arabic generation.
arXiv:2604.08974v2 Announce Type: replace Abstract: Uncertainty quantification techniques measure confidence in language model outputs to support critical applications like hallucination detection an...