arXiv Computation and Language

SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages

SEA-CLIP-Tiny is a compact multilingual text‑vision embedding model designed for Southeast Asian languages, containing fewer than 50 million parameters. It adapts a CLIP‑KD framework with region‑specific data curation and multilingual teacher guidance. Across seven languages, it outperforms other student models, achieving R@1 = 12.9%, R@5 = 31.5%, and R@10 = 42.2%, and surpasses MobileCLIP2 by 12.1 points in R@10 while using 38.4% fewer parameters and lower CPU latency.

arXiv Computation and Language
Aug 25

L\"etzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents

arXiv:2608.21714v1 Announce Type: new Abstract: Recent page-image retrievers such as ColPali have improved retrieval over visually rich documents, yet little is known about how they behave in cross-l...

By Omar El Bachyr, Fred Philippy, Laura Maria Bernardy, Saad Ezzini, Jacques Klein, Tegawende Bissyande
arXiv AI
Aug 28

VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation

The paper introduces Vision-Free Adaptation (VFA), a method that separates multilingual language enhancement from visual alignment in multimodal large language models. VFA fine‑tunes a base LLM on multilingual text to create a multilingual task vector, which is then merged with the vision‑aligned task vector of an existing MLLM. Experiments on five MLLMs and six multilingual benchmarks show consistent gains while preserving multimodal and text‑only performance, and using less than 2% of text data narrows the performance gap to fully multimodal‑trained models.

By Yixia Li, Yaqing Shi, Zhiwen Ruan, Dongdong Zhang, Lingjie Jiang, Shaohan Huang, Yun Chen, Guanhua Chen, Furu Wei
arXiv Computation and Language
Aug 28

Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

BanglaVerse is a new benchmark that evaluates multilingual vision‑language models on Bengali culture, covering nine visual domains and expanding to four languages and five Bangla dialects for a total of about 32,200 artifacts. It includes visual question answering and captioning tasks built from 1,152 manually curated images. Experiments show that models perform worse on dialectal variants and that missing cultural knowledge, rather than visual grounding, is the main bottleneck.

By Nurul Labib Sayeedi, Md. Faiyaz Abdullah Sayeedi, Shubhashis Roy Dipta, Mahbub E Sobhani, Rubaya Tabassum, Ariful Ekraj Hridoy, Mehraj Mahmood, Md. Tarek Hasan, Swakkhar Shatabda
arXiv AI
Jun 12

SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation

arXiv:2606. 13647v1 Announce Type: cross Abstract: We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4$\times$ the depth of existing multilingual benchmark coverage for Slovak.

By Marek \v{S}uppa, Andrej Ridzik, Daniel Hl\'adek, Nat\'alia K\v{n}a\v{z}ekov\'a, Vikt\'oria Ondrejov\'a
arXiv Computation and Language
Sep 1

Evaluating Perspectival Biases in Cross-Modal Retrieval

arXiv:2510.26861v4 Announce Type: replace-cross Abstract: Multimodal retrieval systems are expected to operate in a semantic space, agnostic to the language or cultural origin of the query. In practi...

By Teerapol Saengsukhiran, Peerawat Chomphooyod, Narabodee Rodjananant, Chompakorn Chaksangchaichot, Patawee Prakrankamanant, Witthawin Sripheanpol, Pak Lovichit, Sarana Nutanong, Ekapol Chuangsuwanich