arXiv Computation and Language
The report introduces Nemotron-SEA-LION-v4.8, a family of Southeast Asian language models built on NVIDIA Nemotron 3, featuring 30B-A3B and 120B-A12B variants with both base and post‑trained checkpoints. The models are fine‑tuned on Southeast Asian, reasoning, code, and multilingual parallel datasets, then further refined with supervised fine‑tuning and online on‑policy distillation. On the SEA‑HELM benchmark, the 30B-A3B model raises the overall SEA score from 46.06 to 51.57, while the 120B-A12B model jumps from 49.30 to 63.44, with the largest improvements seen in instruction following, natural language reasoning, and understanding across seven Southeast Asian languages.
SEA-SpeechBench is a large‑scale multitask benchmark for speech understanding in 11 Southeast Asian languages, comprising 97,194 samples across 99 evaluation sets and 597 hours of curated audio. It covers nine tasks in three categories—speech processing, paralinguistic analysis, and a novel temporal understanding dimension—using multilingual prompting in both native SEA languages and English. Evaluation of current models shows significant performance gaps, especially in temporal understanding, emotion recognition, and speech translation, with low‑resource languages lagging behind English by up to 41 percentage points.
By Jingyi Liao, Wenyu Zhang, Zhuohan Liu, Yingxu He, Geyu Lin, Xunlong Zou, Shuo Sun, Syed Ali Redha Alsagoff, Ai Ti Aw
The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leavi...
arXiv:2606.28715v2 Announce Type: replace-cross
Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly und...
By My Chiffon Nguyen, Aulia Adila, Saksorn Ruangtanusak, Kittiphat Leesombatwathana, Vissuta Gunawan Lim, Patomporn Payoungkhamdee, Samuel Cahyawijaya
arXiv:2606.03027v2 Announce Type: replace
Abstract: Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-...
By Peerat Limkonchotiwat, Raymond Ng, Sarana Nutanong, Jian Gang Ngui
arXiv:2609.22586v1 Announce Type: cross
Abstract: Modern audio-language models are no longer judged only on what words they can transcribe, but on whether they can reason over what they hear: recover...
By Harshit Rajgarhia, Asif Shaik, Rachuri Lokesh, Sushanta Kumar Pani, Abhishek Mukherji
The paper investigates how multilingual large language models can be guided to reason more reliably in low- to mid-resource languages by selecting appropriate language modes during inference. Experiments with LLaMA and Qwen models show that using English context can correct errors from non‑English comprehension, but adding redundant bilingual context can cause interference. To balance this trade‑off, the authors propose Reliability‑Aware Adaptive Inference (RAAI), a training‑free test‑time framework that routes prompts based on Expected Calibration Error and gates reasoning with a mid‑layer Risk Index, achieving up to 37.7% accuracy gains and reduced calibration error on low‑resource languages.
By Ekata Mitra, Ameeta Agrawal
LuxIT is a monolingual instruction‑tuning dataset for Luxembourgish, created by synthesizing instruction‑answer pairs from native texts using the DeepSeek‑R1‑0528 model and a quality‑assurance LLM‑as‑judge process. The resulting 227,507 high‑quality pairs were used to fine‑tune 14 LLMs (≤15 B parameters), yielding an average accuracy increase of +5.37 percentage points on standardized Luxembourgish proficiency exams and improvements in macro‑averaged F1 on nine of the fourteen downstream NLP tasks. These findings demonstrate that synthetic monolingual data can effectively enhance LLM performance in low‑resource languages and reveal the complex relationship between exam performance and practical NLP gains.
By Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot
EuroAlpaca presents a task‑preserving localisation pipeline that translates English instruction‑tuning data into 50 European languages while maintaining task‑critical constraints. The method uses field‑wise machine translation or reconstructs task‑equivalent target‑language instances, followed by validation of coherence and consistency. Experiments show that EuroAlpaca improves instruction‑following accuracy by 12.9% over a baseline and outperforms direct translation on ROUGE‑L and F‑BERT metrics.
By Aleix Sant, Jordi Luque, Carlos Escolano
arXiv:2606. 28715v1 Announce Type: cross Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood despite its importance to sovereign AI.
By My Chiffon Nguyen, Aulia Adila, Saksorn Ruangtanusak, Kittiphat Leesombatwathana, Vissuta Gunawan Lim, Patomporn Payoungkhamdee, Samuel Cahyawijaya
arXiv:2410. 07809v2 Announce Type: replace-cross Abstract: Multilingual instruction tuning (MIT) is challenged by the curse of multilinguality, data scarcity, and high computational cost.
By G\"urkan Soykan, G\"ozde G\"ul \c{S}ahin
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y