arXiv:2012. 02110v2 Announce Type: replace-cross Abstract: Pre-trained language models have significantly advanced natural language processing (NLP), especially with the introduction of BERT and its optimized version, RoBERTa.
By Raphael Scheible, Johann Frei, Fabian Thomczyk, Henry He, Patric Tippmann, Jochen Knaus, Victor Jaravine, Frank Kramer, Martin Boeker
The paper introduces a three‑stage training pipeline that builds compact, efficient dense retrievers without requiring ground‑truth relevance labels. Using cross‑lingual alignment, relational knowledge distillation, and contrastive fine‑tuning, the authors develop PolDense (six Polish models ranging from 17 M to 1 B parameters) and EuroDense (a 435 M‑parameter model covering nine European languages). Extensive evaluation on 41 Polish and 150 multilingual tasks shows that PolDense‑1B outperforms larger retrievers up to 9 B parameters, while EuroDense leads in task‑averaged and language‑averaged performance among models below 1 B parameters.
By S{\l}awomir Dadas, Rafa{\l} Po\'swiata, Ma{\l}gorzata Gr\k{e}bowiec, Micha{\l} Pere{\l}kiewicz
arXiv:2606. 31718v1 Announce Type: cross Abstract: Relation extraction (RE) for low-resource languages is typically constrained by the lack of annotated corpora.
By Dragos-Mitrut Vasile, Elena-Simona Apostol, Stefan-Adrian Toma, Adrian Paschke, Ciprian-Octavian Truica
arXiv:2503. 05500v3 Announce Type: replace-cross Abstract: General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models.
By Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, Andr\'e Martins, Ayoub Hammal, Caio Corro, C\'eline Hudelot, Emmanuel Malherbe, Etienne Malaboeuf, Fanny Jourdan, Gabriel Hautreux, Jo\~ao Alves, Kevin El Haddad, Manuel Faysse, Maxime Peyrard, Nuno M. Guerreiro, Patrick Fernandes, Ricardo Rei, Pierre Colombo
Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results....
arXiv:2605. 22641v3 Announce Type: replace-cross Abstract: Detecting Schwartz values in political text is difficult because implicit cues often depend on surrounding arguments and fine-grained distinctions between neighboring values.
By V\'ictor Yeste, Paolo Rosso
LuxIT is a monolingual instruction‑tuning dataset for Luxembourgish, created by synthesizing instruction‑answer pairs from native texts using the DeepSeek‑R1‑0528 model and a quality‑assurance LLM‑as‑judge process. The resulting 227,507 high‑quality pairs were used to fine‑tune 14 LLMs (≤15 B parameters), yielding an average accuracy increase of +5.37 percentage points on standardized Luxembourgish proficiency exams and improvements in macro‑averaged F1 on nine of the fourteen downstream NLP tasks. These findings demonstrate that synthetic monolingual data can effectively enhance LLM performance in low‑resource languages and reveal the complex relationship between exam performance and practical NLP gains.
By Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot
Large language models (LLMs) achieve strong relation extraction (RE), but their computational demands and reliance on proprietary APIs limit deployment in resource-constrained or privacy-sensitive settings. We investigate how far small language models (SLMs) can close this gap across general-domain and literary text.
arXiv:2606. 13647v1 Announce Type: cross Abstract: We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4$\times$ the depth of existing multilingual benchmark coverage for Slovak.
By Marek \v{S}uppa, Andrej Ridzik, Daniel Hl\'adek, Nat\'alia K\v{n}a\v{z}ekov\'a, Vikt\'oria Ondrejov\'a
arXiv:2609.27510v1 Announce Type: cross
Abstract: Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and...
By Kaifeng Tan, Yudong Li, Linlin Shen
R-DEIM Net is a 76‑million‑parameter dual‑expert model designed for paraphrase detection that balances accuracy with computational efficiency. It combines an Interaction Expert, which captures token‑level similarity via multi‑scale 2D convolutions and attention, with a Reasoning Expert that generates human‑readable rationales using a Flan‑T5‑small decoder. On the Quora Question Pairs dataset, the model attains 90.07% accuracy and 90.16% F1‑score, matching strong transformer baselines while producing auxiliary rationales.
By Pushp, Vaibhav Prajapati, Himangshu Sarma
arXiv:2601.14172v4 Announce Type: replace-cross
Abstract: We study neural multi-label classification under severe label imbalance through sentence-level detection of the 19 refined Schwartz human val...
By V\'ictor Yeste, Paolo Rosso