arXiv Machine Learning

Multilingual Coreference Resolution via Cycle-Consistent Machine Translation

arXiv:2606. 05444v1 Announce Type: cross Abstract: Coreference resolution is a core NLP task, having a broad range of downstream applications, e.

arXiv Computation and Language
Sep 10

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

arXiv:2607.00890v2 Announce Type: replace Abstract: Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synth...

By Maximilian Idahl, J\"org Tiedemann, Sampo Pyysalo, David Salinas, Tomasz Galica, Shenbin Qian, Tudor Nicolae Mateiu, Zihao Li, Anna Lokrantz, Fedor Vitiugin, Andr\'e F. T. Martins, Jenna Kanerva, Filip Ginter, Matthias Lindemann, Tim Isbister, Birger Moell, Jonas Lindh, Jan Haji\v{c}, Jenia Jitsev, Andrey Kutuzov, Stephan Oepen, Gema Ram\'irez-S\'anchez
Hugging Face Trending Papers
Jul 20

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens in a response, we naturally guide the model towards generating good tokens, while avoiding the pitfalls that come with directly training the model to generate off-policy tokens.

arXiv Computation and Language
Sep 10

Improving Cross-Lingual Token Representations by Adding a Pinch of SALT

The paper introduces SALT, a lightweight post‑training technique that injects span‑level supervision into existing cross‑lingual sentence encoders to enhance token representations. Evaluated on five multilingual token‑level benchmarks, SALT achieves the best overall results on four tasks, surpassing alternative fine‑tuning methods and competitive encoders. Additionally, SALT improves sentence‑level performance on cross‑lingual retrieval and classification tasks, demonstrating the effectiveness of span‑level supervision for both token and sentence representations.

By Guillem Ram\'irez
arXiv Computation and Language
Sep 25

EnSiTa - A Trilingual Multi-Domain Parallel Dataset and Benchmark for Domain-Specific Machine Translation

EnSiTa is a trilingual multi‑domain parallel dataset and benchmark for English, Sinhala, and Tamil. It contains human post‑edited training data across seven domains and professionally translated test sets for those domains plus an additional one, all produced through a multi‑year, rigorously quality‑controlled process. The authors use EnSiTa to conduct a comprehensive study of domain‑specific machine translation across six language directions, comparing from‑scratch Transformers, pre‑trained models, and decoder‑only LLMs under various training‑data sizes, model scales, and domain settings.

By Surangika Ranathunga, Nisansa de Silva, Aloka Fernando, Kavindu Warnakulasuriya, Isuru Wijesiri, Menan Velayuthan, Charitha Rathnayaka, Thivaharan Varatharajan, Sajeevi Silva, Piumi Kandanaarachchi, Uthayasanker Thayasivam