The paper introduces the Last Translation Benchmark (LTB), a live dataset of human-authored and peer‑reviewed examples—including texts, images, audio, and videos—that are designed to break current state‑of‑the‑art machine translation models. Each example is accompanied by handcrafted verification rules that specify concrete failure cases, providing a reliable and actionable evaluation method. The benchmark aims to overcome the limitations of existing automatic metrics and gold human evaluations, which often lack reproducibility, objectivity, and scalability.
Machine-generated by The Flow from the publisher's headline and feed description
— not written or checked by a human. The full article lives at arXiv Computation and Language.
arXiv:2608. 15932v1 Announce Type: new Abstract: As state-of-the-art machine translation models saturate standard benchmarks, the field needs more challenging evaluations to distinguish between models of varying quality.
By William Kalikman, \v{S}imon Sukup, Michal Te\v{s}nar, Vil\'em Zouhar
As state-of-the-art machine translation models saturate standard benchmarks, the field needs more challenging evaluations to distinguish between models of varying quality. We propose augmenting existi...
arXiv:2609.11399v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used for machine translation, yet their outputs often contain additional text beyond the translation itse...
arXiv:2601.02933v4 Announce Type: replace
Abstract: Human evaluation is the gold standard for multilingual NLP, but is often skipped in practice and substituted with automatic metrics because it is n...
arXiv:2608. 08283v1 Announce Type: cross Abstract: Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation.
By Osvaldo Quinjica, Eric Bennett, Xinchen Yang, Andrew Schonebaum, Marine Carpuat
arXiv:2608. 10812v1 Announce Type: cross Abstract: We study reference-free post-training for multilingual machine translation with open large language models.