arXiv Computer Vision By Guangyu Sun, Shlok Kumar Mishra, Wentao Bao, Robert Zhenheng Yang, Xiao Wang, Xiyuan Wang, Yujunrong Ma, Chen Yuan, Max Xiangjun Fan, Jun Xiao, Jianpeng Cheng

FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

Read the original on arXiv Computer Vision →

FLAT (Flexible‑Length Aligned Transmodal representations) is a joint multimodal pre‑training framework that learns a shared encoder for images and text, producing 1‑D continuous embeddings that can be directly used by downstream generative decoders. By combining contrastive alignment with bidirectional cross‑modal generative objectives, FLAT yields representations that are both discriminative and generative, enabling cross‑modal retrieval and generation with a single pre‑training stage. The model achieves strong performance on T2I generation (GenEval 71.1), image captioning (BLEU‑4 40.5, CIDEr 138.6), and retrieval tasks (Recall@5 86.8/75.8 on MS‑COCO, 98.3/93.6 on Flickr30K), and supports linear interpolation, latent space arithmetic, and zero‑shot composed retrieval.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.