arXiv Machine Learning By Rocker D'Antonio, Thomas Benton Townsend, Dimitrios Michael Manias

Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study

Read the original on arXiv Machine Learning →

The paper investigates how sentence‑specificity scores can guide the selection of revisions in collaborative technical documentation. It compares two predictors—SpeciTeller and a target‑adapted model by Ko et al.—across Wikipedia and three technical corpora, finding that the predictors rank sentences differently and that SpeciTeller can improve direction‑valid selection rates in certain datasets. The study also shows that filtering and token‑length adjustments alter but do not reconcile these ranking differences.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 15

In the Blind: Building Pseudo-References for MT Evaluation

arXiv:2609.13611v1 Announce Type: new Abstract: The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT...

By Diptesh Kanojia, Chi-kiu Lo, Archchana Sindhujan, Samuel Larkin, Greg Hanneman, Alon Lavie
arXiv Computation and Language
4d ago

Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency

Large Language Model judges are commonly used to rank texts via pairwise comparison, with reliability traditionally measured by position bias, transitivity, and pairwise agreement. This paper argues that these proxies are misleading because they are dominated by close‑rank‑gap pairs, which contribute little to the overall ranking, while far‑gap pairs carry the true ranking signal. Experiments on simulations and human‑rated corpora show weak correlation between the proxies and actual ranking accuracy, suggesting judges should be evaluated using rank‑gap‑conditional metrics against human rankings.

By Bruno Brocai, Maria Becker
arXiv AI
Jun 2

Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025

arXiv:2606. 02255v1 Announce Type: cross Abstract: Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and how the annotation process was controlled.

By Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli, Yanran Chen, Christian Greisinger, Lotta Kiefer, Christoph Leiter, Subhadeep Roy, Tewodros Achamaleh, Muhammad Arslan Manzoor, Sebastian Pohl, Yufang Hou, Steffen Eger
arXiv AI
Aug 13

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

arXiv:2608. 12138v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings.

By Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh