arXiv:2508. 05502v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descriptions in low-resource settings.
By Yufei Gao, Jiaying Fei, Nuo Chen, Ruirui Chen, Guohang Yan, Yunshi Lan, Botian Shi
The paper introduces Synth-JDoc, a synthetic Japanese document image dataset created by rendering text with HTML and CSS to produce multi‑column layouts that include both vertical and horizontal writing styles. Images generated by text‑to‑image models are embedded to enhance visual realism, and noise and degradation filters are applied to improve robustness. Experiments show that fine‑tuning Large Vision Language Models on Synth‑JDoc yields superior performance on reading vertically written Japanese text compared to prior synthetic datasets.
By Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara
arXiv:2608. 11002v1 Announce Type: cross Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years.
By Sicheng Zhang, Zhonghao Yan, Binzhu Xie, Shi Qiu, Muzammal Naseer, Naveed Akhtar, Mubarak Shah
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a t...
arXiv:2608.24845v1 Announce Type: cross
Abstract: We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from Commo...
By Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thadd\"aus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Sch\"olkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge
SEA-CLIP-Tiny is a compact multilingual text‑vision embedding model designed for Southeast Asian languages, containing fewer than 50 million parameters. It adapts a CLIP‑KD framework with region‑specific data curation and multilingual teacher guidance. Across seven languages, it outperforms other student models, achieving R@1 = 12.9%, R@5 = 31.5%, and R@10 = 42.2%, and surpasses MobileCLIP2 by 12.1 points in R@10 while using 38.4% fewer parameters and lower CPU latency.
By Puja Ahmad Habibi, Faiz Assabil Firdaus, Ashvanth S, Ekapol Chuangsuwanich, Pume Tuchinda, Peerat Limkonchotiwat
arXiv:2510.26861v4 Announce Type: replace-cross
Abstract: Multimodal retrieval systems are expected to operate in a semantic space, agnostic to the language or cultural origin of the query. In practi...
By Teerapol Saengsukhiran, Peerawat Chomphooyod, Narabodee Rodjananant, Chompakorn Chaksangchaichot, Patawee Prakrankamanant, Witthawin Sripheanpol, Pak Lovichit, Sarana Nutanong, Ekapol Chuangsuwanich
BanglaVerse is a new benchmark that evaluates multilingual vision‑language models on Bengali culture, covering nine visual domains and expanding to four languages and five Bangla dialects for a total of about 32,200 artifacts. It includes visual question answering and captioning tasks built from 1,152 manually curated images. Experiments show that models perform worse on dialectal variants and that missing cultural knowledge, rather than visual grounding, is the main bottleneck.
By Nurul Labib Sayeedi, Md. Faiyaz Abdullah Sayeedi, Shubhashis Roy Dipta, Mahbub E Sobhani, Rubaya Tabassum, Ariful Ekraj Hridoy, Mehraj Mahmood, Md. Tarek Hasan, Swakkhar Shatabda
arXiv:2607. 20241v1 Announce Type: cross Abstract: Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms.
By Yiming Wang, Jiayuan Di
arXiv:2608.20840v1 Announce Type: cross
Abstract: Recent advances in multimodal retrieval have improved the ability to retrieve information from visually rich documents such as PDFs and reports. Howe...
By Yongbin Choi, Yongwoo Song, Mujeen Sung
arXiv:2609.37287v1 Announce Type: cross
Abstract: Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preserva...
By Bo Lv, Mao Zheng, Zheng Li, Fangxu Liu, Mingrui Sun, Tao Chen
arXiv:2602.01161v2 Announce Type: replace
Abstract: The global deployment of large language models (LLMs) has raised concerns about cultural misalignment, yet the linguistic properties of fine-tuning...
By Reem I. Masoud, Chen Feng, Shunta Asano, Saied Alshahrani, Philip Colin Treleaven, Miguel R. D. Rodrigues