A Deepdive into Aya Vision: Advancing the Frontier of Multilingual Multimodality
Related stories
Visual Salamandra: Pushing the Boundaries of Multimodal Understanding
Introducing IDEFICS: An Open Reproduction of State-of-the-art Visual Langage Model
What We are Missing in Multimodal LLM Evaluation?
arXiv:2606. 26348v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can process diverse inputs, e.
SigLIP 2: A better multilingual vision language encoder
Visual Document Retrieval Goes Multilingual
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
arXiv:2608. 11002v1 Announce Type: cross Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years.
MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
arXiv:2508. 05502v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descriptions in low-resource settings.
RMS@CC-MMD 2026: Multimodal Misogyny Detection via Geometric Interaction and Multi-View Consensus
arXiv:2607. 22709v1 Announce Type: cross Abstract: The proliferation of internet memes has introduced new complexities to automated content moderation, particularly in detecting misogyny.
Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models
arXiv:2606. 05531v1 Announce Type: cross Abstract: Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence.