New and improved embedding model
We are excited to announce a new embedding model which is significantly more capable, cost effective, and simpler to use.
We are excited to announce a new embedding model which is significantly more capable, cost effective, and simpler to use.
arXiv:2608. 11343v1 Announce Type: new Abstract: Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning.
arXiv:2606. 14747v1 Announce Type: cross Abstract: Recent advancements have significantly expanded the theoretical context windows of Multimodal Embedding Models (MEMs).
arXiv:2609.37225v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods eith...
The paper introduces Giraffe, a new mapping architecture that converts hidden text token representations into visual embeddings for graphic design tasks. It uses a single [IMG] token per image and two shallow MLP blocks—one for training and one for inference—to compress and expand embeddings, trained with six loss functions. The approach achieves strong performance in both image‑to‑design and text‑to‑design generation while remaining lightweight.
arXiv:2607. 16305v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding.
arXiv:2506. 03096v2 Announce Type: replace-cross Abstract: Contrastive language-image pre-training aligns features of text-image pairs in a common latent space via distinct encoders for each modality.