arXiv:2605.28190v2 Announce Type: replace
Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embeddin...
By Manuel Frank, Haithem Afli
arXiv:2606. 13647v1 Announce Type: cross Abstract: We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4$\times$ the depth of existing multilingual benchmark coverage for Slovak.
By Marek \v{S}uppa, Andrej Ridzik, Daniel Hl\'adek, Nat\'alia K\v{n}a\v{z}ekov\'a, Vikt\'oria Ondrejov\'a
SEA-CLIP-Tiny is a compact multilingual text‑vision embedding model designed for Southeast Asian languages, containing fewer than 50 million parameters. It adapts a CLIP‑KD framework with region‑specific data curation and multilingual teacher guidance. Across seven languages, it outperforms other student models, achieving R@1 = 12.9%, R@5 = 31.5%, and R@10 = 42.2%, and surpasses MobileCLIP2 by 12.1 points in R@10 while using 38.4% fewer parameters and lower CPU latency.
By Puja Ahmad Habibi, Faiz Assabil Firdaus, Ashvanth S, Ekapol Chuangsuwanich, Pume Tuchinda, Peerat Limkonchotiwat
arXiv:2503. 05500v3 Announce Type: replace-cross Abstract: General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models.
By Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, Andr\'e Martins, Ayoub Hammal, Caio Corro, C\'eline Hudelot, Emmanuel Malherbe, Etienne Malaboeuf, Fanny Jourdan, Gabriel Hautreux, Jo\~ao Alves, Kevin El Haddad, Manuel Faysse, Maxime Peyrard, Nuno M. Guerreiro, Patrick Fernandes, Ricardo Rei, Pierre Colombo
arXiv:2608. 05980v1 Announce Type: new Abstract: We investigate whether simple transformations can translate representations across heterogeneous text embedding models.
By Sid Ali Hamideche (Orange Research), Louis Adrien Dufrene (Orange Research), Quentin Lampin (Orange Research), Guillaume Larue (Orange Research)
The paper investigates how to effectively pre‑train language models when the data budget is limited but compute is plentiful. It shows that increasing model size only improves performance up to an optimal point, after which overfitting degrades generalization, and that this optimal size varies with both the data budget and downstream tasks. To overcome the inefficiencies of standard Transformers in this regime, the authors propose recursive Transformers that reuse a shared block across depth and employ factorized embeddings, achieving better results than standard models on 10M–100M word pre‑training budgets and competitive performance with BabyLM Challenge 2025 winners.
By Serdar G\"ulbahar, Lukas Edman, Alexander Fraser
arXiv:2608. 11002v1 Announce Type: cross Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years.
By Sicheng Zhang, Zhonghao Yan, Binzhu Xie, Shi Qiu, Muzammal Naseer, Naveed Akhtar, Mubarak Shah
arXiv:2609.05721v1 Announce Type: new
Abstract: Understanding whether language-model embeddings encode structured real-world information is important for both representation analysis and information...
By Esteban Feuerstein, Victoria Klimkowski, Juan Manuel Ortiz de Zarate, Federico Hern\'an Suaiter
The report introduces Nemotron-SEA-LION-v4.8, a family of Southeast Asian language models built on NVIDIA Nemotron 3, featuring 30B-A3B and 120B-A12B variants with both base and post‑trained checkpoints. The models are fine‑tuned on Southeast Asian, reasoning, code, and multilingual parallel datasets, then further refined with supervised fine‑tuning and online on‑policy distillation. On the SEA‑HELM benchmark, the 30B-A3B model raises the overall SEA score from 46.06 to 51.57, while the 120B-A12B model jumps from 49.30 to 63.44, with the largest improvements seen in instruction following, natural language reasoning, and understanding across seven Southeast Asian languages.
By Ahmed Mohammad Dabeer (David Wang Dawei), Ahn Jeongmi (David Wang Dawei), Anocha Sutaveephamochanon (David Wang Dawei), Antonyrex Sajeban (David Wang Dawei), Aulia Adila (David Wang Dawei), Chan Hok Teng (David Wang Dawei), Adwin (David Wang Dawei), Cheng Zi Yi (David Wang Dawei), Nicholas Zhuang Ziyi (David Wang Dawei), Choa Hsueh Mei Esther (David Wang Dawei), David Ong Tat-Wee (David Wang Dawei), Evelyn Tan Chor Phin (Li Chunren), Heng Cheng Peng (Li Chunren), Jonathan (Li Chunren), Lee Chwan Ren (Li Chunren), Leong Wai Yi (Huang Wenzong, Raymond), Leong Wei Qi (Huang Wenzong, Raymond), Leslie Teo Eng Sipp (Huang Wenzong, Raymond), Liew Rachel (Huang Wenzong, Raymond), Limkonchotiwat Peerat (Huang Wenzong, Raymond), Montalan Jann Railey Estrada (Huang Wenzong, Raymond), Muhammad Ridzuan Bin Mokhtar (Huang Wenzong, Raymond), Nagarajan Karthik (Huang Wenzong, Raymond), Ng Boon Cheong (Huang Wenzong, Raymond), Raymond (Huang Wenzong, Raymond), Ngui Jian Gang (Chen Xiaowei), Nguyen Thanh Ngan (Chen Xiaowei), Tasawong Panuthep (Chen Xiaowei), Pereira Mark Gregory (Chen Xiaowei), Phang Shi Wei Benjamin (Chen Xiaowei), Poon Yip Hung (Chen Xiaowei), Joseph (Chen Xiaowei), Rengarajan Hamsawardhini (Chen Xiaowei), Siow Wei Kang Bryan (Chen Xiaowei), Tai Ngee Chia (Chen Xiaowei), Tan Choon Meng (Chen Xiaowei), Tan Le Min (Chen Xiaowei), Sheryl (Chen Xiaowei), Tan Siao Wei (Chen Xiaowei), Tan Yi Xian, Tee Jun Yun, Teng Kok Wai, Tjhi William Chandra, Tuchinda Pume, Wu Donghang, Yong Xianbin, Yosephine, Zhang Zhou
The study compares an English-only and a bilingual decoder-only model, each 310 M parameters, trained on eight diverse languages while controlling for English exposure, compute, and document overlap. After aligning on shared English vocabulary, the authors find that token embeddings appear similar, but the deeper hidden states used for prediction diverge across models. This hidden‑state mismatch grows through middle transformer layers and persists despite controls, indicating that contextual processing differs between the models.
"whyItMatters":"The findings show that embedding alignment can conceal significant internal representation differences, which is crucial for any downstream work that assumes aligned multilingual models are interchangeable."
By Anjishnu Mukherjee, Ziwei Zhu, Antonios Anastasopoulos