Hugging Face Blog

Train 400x faster Static Embedding Models with Sentence Transformers

arXiv Machine Learning
Aug 28

Squeezing More from Limited Data with Recursive Transformers

The paper investigates how to effectively pre‑train language models when the data budget is limited but compute is plentiful. It shows that increasing model size only improves performance up to an optimal point, after which overfitting degrades generalization, and that this optimal size varies with both the data budget and downstream tasks. To overcome the inefficiencies of standard Transformers in this regime, the authors propose recursive Transformers that reuse a shared block across depth and employ factorized embeddings, achieving better results than standard models on 10M–100M word pre‑training budgets and competitive performance with BabyLM Challenge 2025 winners.

By Serdar G\"ulbahar, Lukas Edman, Alexander Fraser
arXiv AI
Sep 10

Compressing Sequences in the Latent Embedding Space: $K$-Token Merging for Large Language Models

The paper introduces K-Token Merging, a latent-space compression method that merges each contiguous block of K token embeddings into a single embedding using a lightweight encoder. The compressed sequence is then processed by a LoRA-adapted large language model, while generation continues in the original vocabulary. Experiments on tasks such as structural reasoning, sentiment classification, and code editing demonstrate that K-Token Merging achieves up to 75% input length reduction with minimal performance loss, placing it on the Pareto frontier of performance versus compression.

By Zihao Xu, John Harvill, Ziwei Fan, Yizhou Sun, Hao Ding, Hao Wang
arXiv AI
Jun 12

SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation

arXiv:2606. 13647v1 Announce Type: cross Abstract: We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4$\times$ the depth of existing multilingual benchmark coverage for Slovak.

By Marek \v{S}uppa, Andrej Ridzik, Daniel Hl\'adek, Nat\'alia K\v{n}a\v{z}ekov\'a, Vikt\'oria Ondrejov\'a