Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
Related stories
Training and Finetuning Sparse Embedding Models with Sentence Transformers
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers
Train 400x faster Static Embedding Models with Sentence Transformers
Multimodal Embedding & Reranker Models with Sentence Transformers
Train and Fine-Tune Sentence Transformers Models
Training and Finetuning Reranker Models with Sentence Transformers
Train a Sentence Embedding Model with 1B Training Pairs
Introduction to Transformers: an NLP Perspective
arXiv:2311. 17633v2 Announce Type: replace-cross Abstract: Transformers have dominated empirical machine learning models of natural language processing.
Squeezing More from Limited Data with Recursive Transformers
The paper investigates how to effectively pre‑train language models when the data budget is limited but compute is plentiful. It shows that increasing model size only improves performance up to an optimal point, after which overfitting degrades generalization, and that this optimal size varies with both the data budget and downstream tasks. To overcome the inefficiencies of standard Transformers in this regime, the authors propose recursive Transformers that reuse a shared block across depth and employ factorized embeddings, achieving better results than standard models on 10M–100M word pre‑training budgets and competitive performance with BabyLM Challenge 2025 winners.
How Far Do Simple Transformations Translate Across Text Embedding Models?
arXiv:2608. 05980v1 Announce Type: new Abstract: We investigate whether simple transformations can translate representations across heterogeneous text embedding models.