Train 400x faster Static Embedding Models with Sentence Transformers
Related stories
Training and Finetuning Sparse Embedding Models with Sentence Transformers
Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
Train a Sentence Embedding Model with 1B Training Pairs
Train and Fine-Tune Sentence Transformers Models
Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
Multimodal Embedding & Reranker Models with Sentence Transformers
Squeezing More from Limited Data with Recursive Transformers
The paper investigates how to effectively pre‑train language models when the data budget is limited but compute is plentiful. It shows that increasing model size only improves performance up to an optimal point, after which overfitting degrades generalization, and that this optimal size varies with both the data budget and downstream tasks. To overcome the inefficiencies of standard Transformers in this regime, the authors propose recursive Transformers that reuse a shared block across depth and employ factorized embeddings, achieving better results than standard models on 10M–100M word pre‑training budgets and competitive performance with BabyLM Challenge 2025 winners.
Compressing Sequences in the Latent Embedding Space: $K$-Token Merging for Large Language Models
The paper introduces K-Token Merging, a latent-space compression method that merges each contiguous block of K token embeddings into a single embedding using a lightweight encoder. The compressed sequence is then processed by a LoRA-adapted large language model, while generation continues in the original vocabulary. Experiments on tasks such as structural reasoning, sentiment classification, and code editing demonstrate that K-Token Merging achieves up to 75% input length reduction with minimal performance loss, placing it on the Pareto frontier of performance versus compression.
Training and Finetuning Reranker Models with Sentence Transformers
SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation
arXiv:2606. 13647v1 Announce Type: cross Abstract: We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4$\times$ the depth of existing multilingual benchmark coverage for Slovak.