Training and Finetuning Sparse Embedding Models with Sentence Transformers
Related stories
Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
Train 400x faster Static Embedding Models with Sentence Transformers
Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
Train and Fine-Tune Sentence Transformers Models
Training and Finetuning Reranker Models with Sentence Transformers
Multimodal Embedding & Reranker Models with Sentence Transformers
Train a Sentence Embedding Model with 1B Training Pairs
Introduction to Transformers: an NLP Perspective
arXiv:2311. 17633v2 Announce Type: replace-cross Abstract: Transformers have dominated empirical machine learning models of natural language processing.
Squeezing More from Limited Data with Recursive Transformers
The paper investigates how to effectively pre‑train language models when the data budget is limited but compute is plentiful. It shows that increasing model size only improves performance up to an optimal point, after which overfitting degrades generalization, and that this optimal size varies with both the data budget and downstream tasks. To overcome the inefficiencies of standard Transformers in this regime, the authors propose recursive Transformers that reuse a shared block across depth and employ factorized embeddings, achieving better results than standard models on 10M–100M word pre‑training budgets and competitive performance with BabyLM Challenge 2025 winners.
Aligning Sentence Embeddings to Human Concepts via Sparse Autoencoders
arXiv:2607. 00023v1 Announce Type: cross Abstract: Dense sentence embeddings are fundamental to modern Retrieval-Augmented Generation (RAG) systems but suffer from a lack of interpretability due to feature superposition.