Hugging Face Blog

NeoMME: an efficient Multimodal-native and Multilingual Encoder

arXiv AI
Sep 3

NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

NeoMME is a family of 260M and 800M‑parameter multimodal‑native multilingual encoders that process text and raw image patches in a single bidirectional Transformer. Trained from scratch with a masked discrete‑diffusion objective conditioned on visible image patches, NeoMME supports a 16,384‑token context, enabling encoding of up to two 4K UHD images. In downstream tests, NeoMME‑Retriever models outperform all sub‑800M‑parameter baselines on the ViDoRe v3 benchmark and achieve twice the throughput of ColModernVBERT on an NVIDIA L40S, while hierarchical token pooling and asymmetric quantization compress embeddings 255× with minimal loss in retrieval performance.

By Aur\'elien Lac, Tony Wu
arXiv AI
Jun 2

EuroBERT: Scaling Multilingual Encoders for European Languages

arXiv:2503. 05500v3 Announce Type: replace-cross Abstract: General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models.

By Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, Andr\'e Martins, Ayoub Hammal, Caio Corro, C\'eline Hudelot, Emmanuel Malherbe, Etienne Malaboeuf, Fanny Jourdan, Gabriel Hautreux, Jo\~ao Alves, Kevin El Haddad, Manuel Faysse, Maxime Peyrard, Nuno M. Guerreiro, Patrick Fernandes, Ricardo Rei, Pierre Colombo
arXiv Computer Vision
Sep 23

Virtual Encoders in Multimodal Transformers

The paper investigates how multimodal language models can generate perceptual representations without dedicated encoders. It shows that the shared transformer can internally create these representations in its early-to-middle layers, a structure termed a Virtual Encoder. Experiments with linear probing, similarity metrics, and causal analysis reveal that this encoder-like computation emerges even when models receive only perceptual tokens, indicating that perception and language processing can be decoupled within a single architecture.

By Katsuya Ogata, Yuta Nakashima