A Deepdive into Aya Vision: Advancing the Frontier of Multilingual Multimodality
Related stories
Visual Salamandra: Pushing the Boundaries of Multimodal Understanding
PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding
arXiv:2608.21853v1 Announce Type: new Abstract: Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understa...
Introducing IDEFICS: An Open Reproduction of State-of-the-art Visual Langage Model
Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects
BanglaVerse is a new benchmark that evaluates multilingual vision‑language models on Bengali culture, covering nine visual domains and expanding to four languages and five Bangla dialects for a total of about 32,200 artifacts. It includes visual question answering and captioning tasks built from 1,152 manually curated images. Experiments show that models perform worse on dialectal variants and that missing cultural knowledge, rather than visual grounding, is the main bottleneck.
What We are Missing in Multimodal LLM Evaluation?
arXiv:2606. 26348v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can process diverse inputs, e.
SigLIP 2: A better multilingual vision language encoder
NeoMME: an efficient Multimodal-native and Multilingual Encoder
Visual Document Retrieval Goes Multilingual
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
arXiv:2608. 11002v1 Announce Type: cross Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years.
MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
arXiv:2508. 05502v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descriptions in low-resource settings.
RMS@CC-MMD 2026: Multimodal Misogyny Detection via Geometric Interaction and Multi-View Consensus
arXiv:2607. 22709v1 Announce Type: cross Abstract: The proliferation of internet memes has introduced new complexities to automated content moderation, particularly in detecting misogyny.