Ettin Suite: SoTA Paired Encoders and Decoders
Related stories
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
T5Gemma: A new collection of encoder-decoder Gemma models
Introducing T5Gemma, a new collection of encoder-decoder LLMs.
SigLIP 2: A better multilingual vision language encoder
Remote VAEs for decoding with Inference Endpoints 🤗
Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs
arXiv:2606. 03879v1 Announce Type: cross Abstract: As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a prerequisite for principled design.
Universal Assisted Generation: Faster Decoding with Any Assistant Model
SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport
arXiv:2602. 23353v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world.
Separating Representation from Reconstruction Enables Scalable Text Encoders
arXiv:2607. 04011v1 Announce Type: cross Abstract: While decoders have rapidly scaled, encoders have remained largely unchanged since BERT.
EuroBERT: Scaling Multilingual Encoders for European Languages
arXiv:2503. 05500v3 Announce Type: replace-cross Abstract: General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models.
Powerful ASR + diarization + speculative decoding with Hugging Face Inference Endpoints
ARC-Encoder: learning compressed text representations for large language models
arXiv:2510. 20535v2 Announce Type: replace-cross Abstract: Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference costs.