New and improved embedding model
We are excited to announce a new embedding model which is significantly more capable, cost effective, and simpler to use.
Related stories
New embedding models and API updates
Getting Started With Embeddings
🪆 Introduction to Matryoshka Embedding Models
Train 400x faster Static Embedding Models with Sentence Transformers
Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks
The paper introduces TEmBed, a unified benchmark for evaluating tabular embeddings across four representation levels—cell, row, column, and table—using a diverse set of models. It demonstrates that the best model depends on the specific task and representation level, providing practical guidance for selecting embeddings in real-world applications. The study aims to facilitate the development of more general-purpose tabular representation models.
Multi-modal Knowledge Preserving Adapter for Embedding Backward Compatibility
Upgrading embedding models typically requires expensive database re-indexing, as new query embeddings are incompatible with existing database embeddings. While Backward Compatible Training (BCT) mitig...
Build a Domain-Specific Embedding Model in Under a Day
Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Omni-Embed-Mini is a 0.9B‑parameter model that embeds text, speech, audio, images, video, and visually‑rich documents into a single shared cosine space without updating any text‑side parameters. It uses a dense cascaded caption as a teacher signal, allowing the teacher and student to share identical backbone weights and requiring only lightweight projectors and phased LoRA adapters for alignment. The model achieves strong text retrieval performance (49.57 nDCG@10 on MTEB‑v2 BEIR‑8) while extending to five additional modalities and is significantly smaller than other open omni‑modal embedders.
TEmBed-T: A Multi-Dimensional Benchmark for Table-Level Embeddings
arXiv:2607. 24130v1 Announce Type: cross Abstract: Tabular data is the dominant structured-data modality, and learning table representations has become a core research direction.
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retrieval, recommendatio...
Multi-modal Knowledge Preserving Adapter for Embedding Backward Compatibility
The paper introduces the Multi-modal Knowledge Preserving Adapter (MKP-Adapter), an adapter-only approach that enables backward compatible training for multi-modal large language models without updating the backbone. It employs a multi-level preservation loss to maintain embedding geometry and a focal re-weighting strategy to focus on difficult samples. Experiments show strong backward compatibility across image, text, visual document, and video retrieval tasks with minimal latency overhead.