DeepMind Blog

EmbeddingGemma 2: an open, lightweight multimodal embedding model

arXiv AI
Aug 26

Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design

The paper introduces Giraffe, a new mapping architecture that converts hidden text token representations into visual embeddings for graphic design tasks. It uses a single [IMG] token per image and two shallow MLP blocks—one for training and one for inference—to compress and expand embeddings, trained with six loss functions. The approach achieves strong performance in both image‑to‑design and text‑to‑design generation while remaining lightweight.

By Nejla Ghaboosi