Hugging Face Trending Papers

From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets.

arXiv Computer Vision
3d ago

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

PixelUMM is an encoder‑free model that unifies image and video understanding and generation directly in pixel space. It represents images as spatial patches and videos as spatiotemporal tubelets, feeding both through single‑layer linear projections into a shared multimodal backbone. The Mixture‑of‑Transformers architecture blends shared attention with task‑specific parameters, enabling autoregressive text prediction, pixel‑space flow matching, and clean‑pixel video generation, and experiments show competitive performance across tasks while providing design insights for future pixel‑space multimodal models.

By Cong Wei, Xuanchi Ren, Bryan Chu, Weiming Ren, Huan Ling, Jiahui Huang, Laura Leal-Taix\'e, Sanja Fidler, Wenhu Chen, Zian Wang, Jay Zhangjie Wu
arXiv AI
Jun 6

Image Generators are Generalist Vision Learners

arXiv:2604. 20329v3 Announce Type: replace-cross Abstract: Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language understanding and reasoning from generative pretraining.

By Valentin Gabeur, Shangbang Long, Songyou Peng, Paul Voigtlaender, Shuyang Sun, Yanan Bao, Karen Truong, Zhicheng Wang, Wenlei Zhou, Jonathan T. Barron, Kyle Genova, Nithish Kannen, Sherry Ben, Yandong Li, Mandy Guo, Suhas Yogin, Yiming Gu, Huizhong Chen, Oliver Wang, Saining Xie, Howard Zhou, Kaiming He, Thomas Funkhouser, Jean-Baptiste Alayrac, Radu Soricut