arXiv:2601. 22888v4 Announce Type: replace-cross Abstract: More than 80% of the 1.
By Jio Oh, Paul Vicinanza, Thomas Butler, Steven Euijong Whang, Dezhi Hong, Amani Namboori
Dialectal variation remains a major challenge for multilingual language models. Perturbation-based continued pre-training (CPT) has emerged as a promising approach to improving robustness, yet existing work largely evaluates individual perturbation strategies in isolation and provides limited insight into why they work.
arXiv:2606. 03165v1 Announce Type: cross Abstract: The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment).
By Thomas Stephan Juzek, Xiaoyang Ming, Jose A. Hernandez
arXiv:2606. 15521v1 Announce Type: cross Abstract: Tokenization introduces representational redundancy: under a fixed token vocabulary, every byte string admits many valid token encodings, or segmentations, that decode to the same surface string.
By Kanishk Jain, Matthew Day, Tankut Can
arXiv:2509. 26169v2 Announce Type: replace Abstract: Alignment of large language models remains a central challenge in natural language processing.
By Fr\'ed\'eric Berdoz, Luca A. Lanzend\"orfer, Ren\'e Caky, Roger Wattenhofer
arXiv:2607. 19243v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge representations remain disproportionately biased toward high-resource languages.
By Alexander Manev