SigLIP 2: A better multilingual vision language encoder
Related stories
jina-vlm: Small Multilingual Vision Language Model
arXiv:2512. 04032v4 Announce Type: replace-cross Abstract: We present jina-vlm, a token-efficient 2.
Vision Language Models (Better, faster, stronger)
Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models
Accelerating Vision-Language Models: BridgeTower on Habana Gaudi2
A Dive into Vision-Language Models
Let ViT Speak: Generative Language-Image Pre-training
arXiv:2605.00809v3 Announce Type: replace Abstract: In this paper, we present \textbf{Gen}erative \textbf{L}anguage-\textbf{I}mage \textbf{P}re-training (GenLIP), a minimalist generative pretraining...
All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
The paper introduces an all‑in‑one multilingual scene text recognizer called ScriptMoE, which uses a script‑aware mixture‑of‑experts architecture to handle 10 scripts and 229 languages. It is built on a new large‑scale synthetic dataset, TextMuSS‑10M, and evaluated on the TextMuSS‑Bench, achieving 82.06% accuracy—1.31% higher than the best baseline. When integrated into the PP‑OCRv5 pipeline, ScriptMoE raises the end‑to‑end multilingual F1 score from 65.71% to 80.89%, slightly surpassing the best vision‑language model while using far fewer parameters.
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model
arXiv:2606.19100v4 Announce Type: replace Abstract: Large Vision and Language Models (LVLMs) have advanced rapidly, yet European Portuguese (pt-PT) remains systematically underserved by existing open...
LoopVL: Recurrent Visual Intelligence
arXiv:2609.38426v1 Announce Type: cross Abstract: We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-...
Accelerating vision-language models with LFM2.5-VL-DSpark
On the Design Fundamentals of Pixel Text Representation Learning
arXiv:2609.01147v1 Announce Type: cross Abstract: Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders strug...