COMET reports translation quality as a single number, and that number is routinely compared across target languages written in different scripts. Such a comparison assumes Script Invariance: the score...
Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT 2020--2024 shared-task lineage with an extended English--Malayalam resource into \indicqe: $126{,}754$ instances over nine directional pairs, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes.
IndicQE-APE is a consolidated benchmark that unifies quality estimation (QE) and automatic post‑editing (APE) data for nine Indic language pairs, comprising 126,754 instances with multiple aligned labels such as direct assessment, human post‑edit, word‑level OK/BAD tags, and error explanations. The dataset includes a stratified test set across four difficulty axes and supports training and evaluation of six prompted large language models, three COMET metrics, and three APE systems. Experiments reveal that segments with conflicting holistic and token‑level quality signals are consistently ranked lower, while annotator disagreement shows no effect when controlled for score distribution.
whyItMatters":"The benchmark provides a unified resource for training and evaluating QE and APE across Indic languages, enabling consistent comparison of models and metrics on a shared dataset."
By Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Or\u{a}san, Chrysoula Zerva, Ricardo Rei, Fr\'ed\'eric Blain, Andr\'e F. T. Martins, Marco Turchi, Matteo Negri, Anoop Kunchukuttan, Mitesh M. Khapra, Pushpak Bhattacharyya
arXiv:2608.16344v3 Announce Type: replace
Abstract: Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and...
By Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Or\u{a}san, Chrysoula Zerva, Ricardo Rei, Fr\'ed\'eric Blain, Andr\'e F. T. Martins, Marco Turchi, Matteo Negri, Anoop Kunchukuttan, Mitesh M. Khapra, Pushpak Bhattacharyya
arXiv:2608. 16643v1 Announce Type: cross Abstract: Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation.
By Yifan Zhang, Rahmatollah Beheshti
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y
SyntaxBench is a diagnostic benchmark and statistical evaluation framework for character‑level reasoning in large language models, comprising five core tasks—character counting, letter containment, palindrome detection, edit distance, and longest‑string selection—and a harder substring‑extraction stress test called index_to_span. The benchmark uses paired English and random‑string inputs, zero‑, one‑, and four‑shot prompts, and evaluates models from 2B to 32B parameters across multiple reasoning modes. It reports a wide range of metrics, including exact‑match and relaxed accuracy, Cohen’s kappa, McNemar tests, bootstrap confidence intervals, Kendall’s tau, class‑conditional metrics, tokenization analysis, and multiple‑comparison‑corrected tests.
By Mohsen Larni (Department of Computer Science, University of Nevada, Las Vegas), Sobhan Ebrahimi Azar (Department of Computer Science, University of Nevada, Las Vegas), Pouyan Nahed (Department of Computer Science, University of Nevada, Las Vegas), Kazem Taghva (Department of Computer Science, University of Nevada, Las Vegas)
Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure does not treat languages equally. One underexamined source of disparity is tokenization: semantically equivalent content can require substantially different token counts across languages, affecting API cost, latency, and usable context length before a model is invoked.
Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation. Error-detection benchmarks are typically constructed by injecting errors into notes, such that each erroneous note has a natural counterpart.
arXiv:2507.23248v2 Announce Type: replace-cross
Abstract: Bengali is spoken by more than 230 million people, yet no standardized instrument evaluates large language models (LLMs) on Bengali across th...
By Shimanto Bhowmik, Tawsif Tashwar Dipto, Md Sazzad Islam, Sheryl Hsu, Tahsin Reasat
arXiv:2609.16340v1 Announce Type: cross
Abstract: Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post-edits of an older system'...
By Rohit Dhaipule, Sukhdeep Singh Kharbanda, Prasanth Bathala, Pradyumna Lanka, Anubhav Shrimal
arXiv:2609.14963v1 Announce Type: new
Abstract: As large language models become capable translators of classical texts, a key challenge is deciding which outputs need expert review when no human refe...
By M\'at\'e Metzger