arXiv:2609.37883v1 Announce Type: new
Abstract: Pre-trained language models (PLMs) with encoder-based architectures have shown impressive capabilities in zero-shot cross-lingual transfer for various...
By Lalita Lowphansirikul, Attapol Rutherford, Jian Gang Ngui, Sarana Nutanong, Peerat Limkonchotiwat
arXiv:2606. 12708v1 Announce Type: cross Abstract: Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP.
By Happy Buzaaba, Cheikh Mouhamadou Bamba Dione, David Ifeoluwa Adelani, Sylvain Kahane, Kim Gerdes, Bruno Guillaume, Kevin Guan, Aremu Anuoluwapo, Naome A. Etori, Shamsuddeen Hassan Muhammad, Utitofon Inyang, Peter Nabende, David Sabiiti Bamutura, Andiswa Bukula, Chinedu Uchechukwu, Rooweither Mabuya, Idris Akinade, Christiane Fellbaum
The paper introduces AraSEG, a new Arabic sentence segmentation corpus covering eight genres and diverse punctuation and document structures. Experiments using AraSEG evaluate large language models, lightweight encoders, and dependency parser-based models, revealing that lightweight encoders and parser-based models outperform LLMs under the most challenging conditions. The study also shows that increasing training data size and genre diversity eventually saturates performance, that cross‑genre generalization remains difficult, and that accurate sentence segmentation significantly improves downstream dependency parsing.
By Mohammed Elkholy, Khalid N. Elmadani, Nizar Habash, Bashar Alhafni
arXiv:2608.23448v1 Announce Type: new
Abstract: This paper presents new Cantonese ParGram resources and evaluates LLMs for knowledge-driven grammar engineering within a controlled experimental paradi...
By Chit-Fung Lam
arXiv:2608. 06179v1 Announce Type: new Abstract: Preference optimisation has proven effective for improving large language models but typically relies on costly human preference annotations.
By Hoda Fakharzadehjahromy, Emil Wiman, Andreas Bueff, Hafsteinn Einarsson, Fredrik Heintz
The paper presents a pipeline that leverages large language models to extract grammatical rules, example sentences, and lexicons from descriptive grammar books, producing synthetic parallel corpora for fine‑tuning machine translation models. Evaluated on three low‑resource languages—Kalamang, Tuatschin, and Mandan—the synthetic data improves translation quality over seed‑data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, achieving up to +8.8 ChrF++ gains. A factorial study across 96 configurations identifies which combinations of target part‑of‑speech, retrieval granularity, and sample volume drive performance gains and where they fail, demonstrating that static linguistic documentation can be repurposed for practical translation tools for severely under‑resourced languages.
By Varun Ghat Ravikumar, Sina Ahmadi, Lena J\"ager, Rico Sennrich
arXiv:2608. 10137v1 Announce Type: cross Abstract: Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step.
By I\c{s}{\i}l \"Ozg\"u, Yaoxuan Wu, Guy Van den Broeck, Miryung Kim
Preference optimisation has proven effective for improving large language models but typically relies on costly human preference annotations. Extending these methods to morphologically rich, low-resource languages remains challenging because such annotations are scarce.
arXiv:2606. 18856v1 Announce Type: cross Abstract: Sequence labelling, a core task of Natural Language Processing (NLP), consists in assigning each token of an input sentence a label.
By Nicolas Floquet, Joseph Le Roux, Nadi Tomeh
arXiv:2606. 02991v1 Announce Type: cross Abstract: We introduce TypewriterLM, a 7.
By Xiaoxi Luo, Zachary Shinnick, Niclas Griesshaber, Yixuan Wang, Junchi Yu, Freda Shi, Philip Torr, Yao Lu
arXiv:2603.01243v3 Announce Type: replace
Abstract: Large language models (LLMs) are powerful tools that have found applications beyond human-machine interfaces and chatbots. Beside free-form generat...
By Ayoub Hammal, Pierre Zweigenbaum, Caio Corro
The paper proposes a modular tokenizer framework for multilingual large language models, allowing the creation of language‑specific subtokenizers that match monolingual compression quality. It introduces a pretraining strategy that samples these subtokenizers to limit predictions to relevant vocabularies, enabling efficient training and inference. This approach reduces memory usage and speeds up inference without compromising performance.
By Franck Signe, Hippolyte Pilchen, Fran\c{c}ois Yvon, \'Edouard Grave