arXiv AI

Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web

The paper introduces a pipeline for creating a high‑quality European Portuguese (PT‑PT) web corpus, drawing from 411 TB of raw data from Arquivo.pt. It adds a novel post‑scraping step that removes boilerplate and duplicate lines before filtering, boosting the final document yield by 19.04%. The pipeline also incorporates language identification, weighted fuzzy deduplication, and neural quality classification to produce a clean, representative dataset suitable for large‑language‑model pre‑training.

arXiv AI
Jul 28

CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data

arXiv:2607. 22662v1 Announce Type: new Abstract: Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance.

By Peiguang Li, Yongwei Zhou, Juncheng Diao, Yuchun Fan, Jian Yang, Jianxiao Yang, Zhongda Su, Shuguang Jiao, Xiao Wei, Zhiye Zou, Gan Dong, Zhizhao Zeng, Rongxiang Weng, Jingang Wang, Xunliang Cai
arXiv Computation and Language
Sep 10

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

arXiv:2607.00890v2 Announce Type: replace Abstract: Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synth...

By Maximilian Idahl, J\"org Tiedemann, Sampo Pyysalo, David Salinas, Tomasz Galica, Shenbin Qian, Tudor Nicolae Mateiu, Zihao Li, Anna Lokrantz, Fedor Vitiugin, Andr\'e F. T. Martins, Jenna Kanerva, Filip Ginter, Matthias Lindemann, Tim Isbister, Birger Moell, Jonas Lindh, Jan Haji\v{c}, Jenia Jitsev, Andrey Kutuzov, Stephan Oepen, Gema Ram\'irez-S\'anchez
arXiv AI
3d ago

Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning

The paper introduces the problem of cross‑lingual loopholes in large language model (LLM) unlearning, where forgetting a fact in one language can leave it accessible in others. It presents a new 174‑language benchmark, the Cross‑Lingual Unlearning Tensor, and proposes COVER, a method that selects a subset of source languages to maximize unlearning coverage under a language budget. Experiments show COVER reduces residual knowledge by 7.8–27.3% compared to uniform selection and works on both synthetic and real low‑resource news data.

By Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav
arXiv Computation and Language
Sep 11

TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs

TransClean introduces a benchmark for identifying and removing translation noise—unwanted text such as language labels, explanations, or bilingual repetitions—from large language model (LLM) outputs. The authors analyzed 790,000 translations from 12 LLMs across 22 language pairs, cataloguing 12 common noise patterns and creating 9,900 noisy‑clean pairs (8,800 synthetic, 1,100 authentic). They evaluated two extraction methods—a span‑based approach using quality estimation models and an LLM‑prompted method—demonstrating the first systematic framework to assess and improve translation cleanliness.

By Shenbin Qian, Yves Scherrer
arXiv Computation and Language
Sep 14

Parameter-Efficient Retrievers for Polish and European Languages

The paper introduces a three‑stage training pipeline that builds compact, efficient dense retrievers without requiring ground‑truth relevance labels. Using cross‑lingual alignment, relational knowledge distillation, and contrastive fine‑tuning, the authors develop PolDense (six Polish models ranging from 17 M to 1 B parameters) and EuroDense (a 435 M‑parameter model covering nine European languages). Extensive evaluation on 41 Polish and 150 multilingual tasks shows that PolDense‑1B outperforms larger retrievers up to 9 B parameters, while EuroDense leads in task‑averaged and language‑averaged performance among models below 1 B parameters.

By S{\l}awomir Dadas, Rafa{\l} Po\'swiata, Ma{\l}gorzata Gr\k{e}bowiec, Micha{\l} Pere{\l}kiewicz
arXiv Machine Learning
Sep 17

TACTICS: Taxonomy-Aware Intelligent Corpus Sampling for Machine Translation

TACTICS is a method for selecting evaluation samples in machine translation that explicitly optimizes for coverage of rare linguistic categories, document-level coherence, and distributional fidelity to the full corpus. It builds a hierarchical taxonomy from a locale style guide, classifies segments, and chooses a fixed-budget subset that better represents the full range of phenomena a system must handle. Compared to random, lexical, or embedding-based selection, TACTICS improves coverage of rare categories and yields more accurate system rankings with fewer segments.

By Prasanth Bathala, Anubhav Shrimal, Sukhdeep Singh Kharbhanda, Pradyumna Lanka, Rohit Dhaipule