arXiv:2609.00998v1 Announce Type: cross
Abstract: Word Sense Disambiguation has advanced rapidly for English and a handful of well-resourced modern languages, but it continues to assume the existence...
By Michele Ciletti
Word Sense Disambiguation has advanced rapidly for English and a handful of well-resourced modern languages, but it continues to assume the existence of a sense inventory and a word-to-sense mapping i...
arXiv:2607.00171v2 Announce Type: replace
Abstract: Text embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge. Current benchmarks are static, cover on...
By Andrianos Michail, Stylianos Psychias, Michelle Wastl, Simon Clematide, Rico Sennrich, Juri Opitz
The paper explores how to improve literary machine translation by using datasets that contain multiple valid translations of the same source text. It introduces a filtering framework that selects source texts whose references show meaningful variation while staying faithful, based on semantic similarity. Experiments show that fine‑tuning on medium to high similarity data outperforms low similarity data, and that using only this filtered subset can match or exceed performance achieved with the full unfiltered set. Additionally, the study compares synthetic translations generated by large language models with human expert translations, finding that fine‑tuning on human expert data yields better results in both automatic metrics and human evaluations, underscoring the continued importance of expert translations for literary MT.
By Si Wu, John Wieting, David A. Smith
arXiv:2606.13218v2 Announce Type: replace
Abstract: Arabic and Hebrew, as closely related Semitic languages, share many words with similar surface forms, including true cognates, false friends, and m...
By Junhong Liang, Noor Abo Mokh, Bashar Alhafni
The paper investigates how to fairly compare language models across languages, noting that current evaluation methods vary widely and lack empirical validation. By training controlled monolingual models on parallel data and testing multilingual LLMs, the authors find that many normalized metrics suffer from biases due to tokenization, encoding, and orthographic differences. Instead, they recommend using sentence‑level negative log‑likelihood over semantically equivalent sequences for more reliable cross‑lingual comparisons.
By Xiulin Yang, Ethan Gotlieb Wilcox, Catherine Arnett