Multimodal Embedding & Reranker Models with Sentence Transformers
Related stories
Training and Finetuning Reranker Models with Sentence Transformers
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
Training and Finetuning Embedding Models with Sentence Transformers
Training and Finetuning Sparse Embedding Models with Sentence Transformers
Train 400x faster Static Embedding Models with Sentence Transformers
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval
arXiv:2608. 11343v1 Announce Type: new Abstract: Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning.
Train and Fine-Tune Sentence Transformers Models
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
arXiv:2506. 03096v2 Announce Type: replace-cross Abstract: Contrastive language-image pre-training aligns features of text-image pairs in a common latent space via distinct encoders for each modality.
Multimodal Representation Alignment for Cross-modal Information Retrieval
arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.
MMLongEmbed: Benchmarking Multimodal Embedding Models in Long-Context Scenarios
arXiv:2606. 14747v1 Announce Type: cross Abstract: Recent advancements have significantly expanded the theoretical context windows of Multimodal Embedding Models (MEMs).
Multimodal Generative Engine Optimization: Rank Manipulation for Vision-Language Model Rankers
arXiv:2601. 12263v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) integrate visual and textual knowledge into unified representations that increasingly underpin modern retrieval and recommendation systems.