Hugging Face Blog

NVIDIA Releases 6 Million Multi-Lingual Reasoning Dataset

arXiv Computation and Language
Sep 17

SEA-LION-v4.8: A Technical Report

The report introduces Nemotron-SEA-LION-v4.8, a family of Southeast Asian language models built on NVIDIA Nemotron 3, featuring 30B-A3B and 120B-A12B variants with both base and post‑trained checkpoints. The models are fine‑tuned on Southeast Asian, reasoning, code, and multilingual parallel datasets, then further refined with supervised fine‑tuning and online on‑policy distillation. On the SEA‑HELM benchmark, the 30B-A3B model raises the overall SEA score from 46.06 to 51.57, while the 120B-A12B model jumps from 49.30 to 63.44, with the largest improvements seen in instruction following, natural language reasoning, and understanding across seven Southeast Asian languages.

By Ahmed Mohammad Dabeer (David Wang Dawei), Ahn Jeongmi (David Wang Dawei), Anocha Sutaveephamochanon (David Wang Dawei), Antonyrex Sajeban (David Wang Dawei), Aulia Adila (David Wang Dawei), Chan Hok Teng (David Wang Dawei), Adwin (David Wang Dawei), Cheng Zi Yi (David Wang Dawei), Nicholas Zhuang Ziyi (David Wang Dawei), Choa Hsueh Mei Esther (David Wang Dawei), David Ong Tat-Wee (David Wang Dawei), Evelyn Tan Chor Phin (Li Chunren), Heng Cheng Peng (Li Chunren), Jonathan (Li Chunren), Lee Chwan Ren (Li Chunren), Leong Wai Yi (Huang Wenzong, Raymond), Leong Wei Qi (Huang Wenzong, Raymond), Leslie Teo Eng Sipp (Huang Wenzong, Raymond), Liew Rachel (Huang Wenzong, Raymond), Limkonchotiwat Peerat (Huang Wenzong, Raymond), Montalan Jann Railey Estrada (Huang Wenzong, Raymond), Muhammad Ridzuan Bin Mokhtar (Huang Wenzong, Raymond), Nagarajan Karthik (Huang Wenzong, Raymond), Ng Boon Cheong (Huang Wenzong, Raymond), Raymond (Huang Wenzong, Raymond), Ngui Jian Gang (Chen Xiaowei), Nguyen Thanh Ngan (Chen Xiaowei), Tasawong Panuthep (Chen Xiaowei), Pereira Mark Gregory (Chen Xiaowei), Phang Shi Wei Benjamin (Chen Xiaowei), Poon Yip Hung (Chen Xiaowei), Joseph (Chen Xiaowei), Rengarajan Hamsawardhini (Chen Xiaowei), Siow Wei Kang Bryan (Chen Xiaowei), Tai Ngee Chia (Chen Xiaowei), Tan Choon Meng (Chen Xiaowei), Tan Le Min (Chen Xiaowei), Sheryl (Chen Xiaowei), Tan Siao Wei (Chen Xiaowei), Tan Yi Xian, Tee Jun Yun, Teng Kok Wai, Tjhi William Chandra, Tuchinda Pume, Wu Donghang, Yong Xianbin, Yosephine, Zhang Zhou
arXiv Computation and Language
Aug 27

Rethinking the Multilingual Reasoning Gap with Layer Swap

The study investigates the performance gap between native-language reasoning and English-pivoted reasoning in large language models. By creating extensive multilingual reasoning datasets and fine‑tuning specialists on Qwen/Qwen3-8B-Base, the authors find that the native reasoning gap is much smaller (1.9–3.5%) than previously reported. They analyze weight‑space changes, discover a language‑agnostic reasoning core in the middle layers, and propose a Layer Swap technique that transfers these mid‑layer updates from an English specialist to native specialists, effectively closing most of the gap while maintaining native chain‑of‑thought output.

By Maxence Lasbordes, Am\'elie Chatelain, Djam\'e Seddah
arXiv AI
Aug 19

BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

BEAR-Bench is a bilingual benchmark for multimodal large language models, featuring 1,000 human‑annotated questions derived from text‑rich business and scientific documents in English and Russian. It evaluates 16 MLLMs, including Gemini 3.1 Pro and Qwen3.5‑397B, revealing significant performance gaps even for the strongest systems. The benchmark also serves to compare hallucination‑detection methods by analyzing model failures on these complex documents.

By Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov, Alexey Zaytsev
arXiv AI
Jun 12

ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages

arXiv:2606. 13572v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in multilingual and low-resource scenarios.

By Tanmoy Kanti Halder, Akash Ghosh, Subhadip Baidya, Arijit Roy, Sriparna Saha
arXiv AI
Sep 7

Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

The paper surveys efficient reasoning in large language models, contrasting fast intuitive (System 1) and slow deep (System 2) reasoning. It analyzes why System 2 is computationally costly yet more accurate, and why System 1 is efficient but less effective. The survey covers causes of inefficiency, patterns of reasoning behavior, and potential solutions to balance performance and computational budgets, offering actionable insights and an open‑source repository for ongoing research.

By Rui Wang, Hongru Wang, Boyang Xue, Jianhui Pang, Shudong Liu, Yi Chen, Jiahao Qiu, Derek Fai Wong, Heng Ji, Kam-Fai Wong
arXiv Computation and Language
Aug 31

Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts

The paper investigates why large reasoning language models struggle to transfer parametric knowledge across different scripts. Through observational data and regression analysis on ECLeKTic and MultiLoKo datasets, the authors find that script mismatch—not language family—is the main predictor of transfer failure when controlling for model capability and question difficulty. By providing key entities in the source language and training models to reason about transliteration ambiguities, they demonstrate a reduction in the cross‑script transfer gap, suggesting that post‑training improvements can enhance cross‑lingual knowledge transfer.

By Lucas Bandarkar, Alan Ansell, Trevor Cohn