FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
FLAT (Flexible‑Length Aligned Transmodal representations) is a joint multimodal pre‑training framework that learns a shared encoder for images and text, producing 1‑D continuous embeddings that can be directly used by downstream generative decoders. By combining contrastive alignment with bidirectional cross‑modal generative objectives, FLAT yields representations that are both discriminative and generative, enabling cross‑modal retrieval and generation with a single pre‑training stage. The model achieves strong performance on T2I generation (GenEval 71.1), image captioning (BLEU‑4 40.5, CIDEr 138.6), and retrieval tasks (Recall@5 86.8/75.8 on MS‑COCO, 98.3/93.6 on Flickr30K), and supports linear interpolation, latent space arithmetic, and zero‑shot composed retrieval.
arXiv:2603. 22282v2 Announce Type: replace-cross Abstract: We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single architecture.
arXiv:2608. 12987v1 Announce Type: cross Abstract: Generative information retrieval (GIR) has emerged as a compelling alternative to the conventional index-retrieve-then-rank retrieval pipeline by training a generator to produce the identifiers of relevant items directly.
arXiv:2508. 12466v2 Announce Type: replace-cross Abstract: Traditional multimodal learning approaches rely on alignment pre-training to bridge vision and language modalities, typically by projecting visual features into discrete text token spaces using large-scale image--text data.
arXiv:2509. 07295v4 Announce Type: replace-cross Abstract: Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture.
arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.