New and improved embedding model
We are excited to announce a new embedding model which is significantly more capable, cost effective, and simpler to use.
We are excited to announce a new embedding model which is significantly more capable, cost effective, and simpler to use.
Upgrading embedding models typically requires expensive database re-indexing, as new query embeddings are incompatible with existing database embeddings. While Backward Compatible Training (BCT) mitig...
The paper introduces the Multi-modal Knowledge Preserving Adapter (MKP-Adapter), an adapter-only approach that enables backward compatible training for multi-modal large language models without updating the backbone. It employs a multi-level preservation loss to maintain embedding geometry and a focal re-weighting strategy to focus on difficult samples. Experiments show strong backward compatibility across image, text, visual document, and video retrieval tasks with minimal latency overhead.
We are introducing embeddings, a new endpoint in the OpenAI API that makes it easy to perform natural language and code tasks like semantic search, clustering, topic modeling, and classification.
arXiv:2508. 21290v2 Announce Type: replace-cross Abstract: jina-code-embeddings is a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify semantically similar code snippets across programming languages.
arXiv:2609.05721v1 Announce Type: new Abstract: Understanding whether language-model embeddings encode structured real-world information is important for both representation analysis and information...