arXiv Computation and LanguageBy Ahmed Mohammad Dabeer (David Wang Dawei), Ahn Jeongmi (David Wang Dawei), Anocha Sutaveephamochanon (David Wang Dawei), Antonyrex Sajeban (David Wang Dawei), Aulia Adila (David Wang Dawei), Chan Hok Teng (David Wang Dawei), Adwin (David Wang Dawei), Cheng Zi Yi (David Wang Dawei), Nicholas Zhuang Ziyi (David Wang Dawei), Choa Hsueh Mei Esther (David Wang Dawei), David Ong Tat-Wee (David Wang Dawei), Evelyn Tan Chor Phin (Li Chunren), Heng Cheng Peng (Li Chunren), Jonathan (Li Chunren), Lee Chwan Ren (Li Chunren), Leong Wai Yi (Huang Wenzong, Raymond), Leong Wei Qi (Huang Wenzong, Raymond), Leslie Teo Eng Sipp (Huang Wenzong, Raymond), Liew Rachel (Huang Wenzong, Raymond), Limkonchotiwat Peerat (Huang Wenzong, Raymond), Montalan Jann Railey Estrada (Huang Wenzong, Raymond), Muhammad Ridzuan Bin Mokhtar (Huang Wenzong, Raymond), Nagarajan Karthik (Huang Wenzong, Raymond), Ng Boon Cheong (Huang Wenzong, Raymond), Raymond (Huang Wenzong, Raymond), Ngui Jian Gang (Chen Xiaowei), Nguyen Thanh Ngan (Chen Xiaowei), Tasawong Panuthep (Chen Xiaowei), Pereira Mark Gregory (Chen Xiaowei), Phang Shi Wei Benjamin (Chen Xiaowei), Poon Yip Hung (Chen Xiaowei), Joseph (Chen Xiaowei), Rengarajan Hamsawardhini (Chen Xiaowei), Siow Wei Kang Bryan (Chen Xiaowei), Tai Ngee Chia (Chen Xiaowei), Tan Choon Meng (Chen Xiaowei), Tan Le Min (Chen Xiaowei), Sheryl (Chen Xiaowei), Tan Siao Wei (Chen Xiaowei), Tan Yi Xian, Tee Jun Yun, Teng Kok Wai, Tjhi William Chandra, Tuchinda Pume, Wu Donghang, Yong Xianbin, Yosephine, Zhang Zhou
The report introduces Nemotron-SEA-LION-v4.8, a family of Southeast Asian language models built on NVIDIA Nemotron 3, featuring 30B-A3B and 120B-A12B variants with both base and post‑trained checkpoints. The models are fine‑tuned on Southeast Asian, reasoning, code, and multilingual parallel datasets, then further refined with supervised fine‑tuning and online on‑policy distillation. On the SEA‑HELM benchmark, the 30B-A3B model raises the overall SEA score from 46.06 to 51.57, while the 120B-A12B model jumps from 49.30 to 63.44, with the largest improvements seen in instruction following, natural language reasoning, and understanding across seven Southeast Asian languages.
Machine-generated by The Flow from the publisher's headline and feed description
— not written or checked by a human. The full article lives at arXiv Computation and Language.
SEA-SpeechBench is a large‑scale multitask benchmark for speech understanding in 11 Southeast Asian languages, comprising 97,194 samples across 99 evaluation sets and 597 hours of curated audio. It covers nine tasks in three categories—speech processing, paralinguistic analysis, and a novel temporal understanding dimension—using multilingual prompting in both native SEA languages and English. Evaluation of current models shows significant performance gaps, especially in temporal understanding, emotion recognition, and speech translation, with low‑resource languages lagging behind English by up to 41 percentage points.
By Jingyi Liao, Wenyu Zhang, Zhuohan Liu, Yingxu He, Geyu Lin, Xunlong Zou, Shuo Sun, Syed Ali Redha Alsagoff, Ai Ti Aw
The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leavi...
arXiv:2606.28715v2 Announce Type: replace-cross
Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly und...
By My Chiffon Nguyen, Aulia Adila, Saksorn Ruangtanusak, Kittiphat Leesombatwathana, Vissuta Gunawan Lim, Patomporn Payoungkhamdee, Samuel Cahyawijaya
arXiv:2606.03027v2 Announce Type: replace
Abstract: Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-...
By Peerat Limkonchotiwat, Raymond Ng, Sarana Nutanong, Jian Gang Ngui
arXiv:2609.22586v1 Announce Type: cross
Abstract: Modern audio-language models are no longer judged only on what words they can transcribe, but on whether they can reason over what they hear: recover...