jina-vlm: Small Multilingual Vision Language Model
arXiv:2512. 04032v4 Announce Type: replace-cross Abstract: We present jina-vlm, a token-efficient 2.
arXiv:2512. 04032v4 Announce Type: replace-cross Abstract: We present jina-vlm, a token-efficient 2.
Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latency.
arXiv:2503. 05500v3 Announce Type: replace-cross Abstract: General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models.
arXiv:2608. 12333v1 Announce Type: cross Abstract: Vision-language models must associate visual entities with textual attributes.
arXiv:2608. 19726v1 Announce Type: cross Abstract: The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder.