arXiv:2509. 24900v2 Announce Type: replace-cross Abstract: The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data.
By Zhihong Chen, Xuehai Bai, Yang Shi, Chaoyou Fu, Huanyu Zhang, Haotian Wang, Xiaoyan Sun, Zhang Zhang, Liang Wang, Yuanxing Zhang, Pengfei Wan, Yi-Fan Zhang
arXiv:2606. 31603v1 Announce Type: cross Abstract: Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.
By Nikolai R\"ohrich, Julian Glei{\ss}ner, Ahmed H. A. Ibrahim, Silvan Mertes, Tobias Huber
RefVideo-6M is a new large-scale reference-guided editing dataset that includes 5 million video editing samples and 1 million image editing samples, each paired with about 6 million visual references. The dataset is constructed to avoid artifacts by using real, artifact‑free videos as targets and filtering input conditions with multiple editing experts, thereby providing reliable supervision. It enables models to learn fine‑grained visual correspondence beyond text‑only instructions and supports the training of a reference‑guided video editing model, Ref‑MoT, which shows improved visual quality, controllability, and reference consistency.
By Bojia Zi, Xiaoyan Yang, Yu Zhou, Ruijie Sun, Lihan Zhang, Bin Liang, Kam-Fai Wong, Haibin Huang, Chi Zhang, Xuelong Li
arXiv:2607.18227v2 Announce Type: replace
Abstract: In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and imag...
By Dingyun Zhang, Lixue Gong, Wei Liu
arXiv:2410. 21361v2 Announce Type: replace-cross Abstract: Domain adaptation has been extensively investigated in computer vision but still requires access to target data at the training time, which might be difficult to obtain in real-world autonomous driving scenarios, especially under rare or adverse conditions.
By Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick P\'erez, Raoul de Charette
The paper introduces a post‑training approach that enables a single inference process to transition from text reasoning to image synthesis, eliminating the need for explicit modality switching. Using the 14B BAGEL model, the authors demonstrate that targeted post‑training data and reward‑weighted training improve multimodal image generation across four independent T2I benchmarks. The study highlights the benefits of joint text‑image generation and strategic data selection for enhancing T2I performance.
By Jiahui Chen, Philippe Hansen-Estruch, Xiaochuang Han, Yushi Hu, Emily Dinan, Amita Kamath, Michal Drozdzal, Reyhane Askari-Hemmat, Luke Zettlemoyer, Marjan Ghazvininejad
The paper introduces a capability‑centric data infrastructure for generalist image generation, integrating task‑specific supervision with a curriculum that aligns with the dependencies among generative capabilities. It employs three interoperable data engines—text‑image grounding, inter‑image transformation, and image‑knowledge association—alongside caption experts to harmonize text‑to‑image and editing supervision. The system curates massive corpora (440M T2I images, 120M editing pairs, 27M image‑entity pairs) and trains multimodal diffusion models (3B and 6B parameters) from scratch, achieving broad visual coverage and versatile rendering as shown by CPI‑Bench and qualitative tests.
By Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen
ProCAP introduces a probabilistic cross-attentive prompt learning framework for vision-language models like CLIP, enabling improved cross-modal interaction without updating the backbone. It jointly learns visual and textual prompt tokens, linking them via stacked bidirectional multi-head cross-attention to refine each branch across prompt depth. The method incorporates Gaussian parameterization of prompt tokens, lightweight KL and L2 regularization, and a compact symmetric InfoNCE head to align image features with class-level text representations, achieving strong few-shot base-to-novel performance and competitive transfer results across multiple datasets and benchmarks.
By Hiwa Azeez Abbas, Fatemeh Daneshfar, Moloud Abdar
ARGenSeg introduces an autoregressive generation-based approach for image segmentation that integrates seamlessly with multimodal large language models (MLLMs). Unlike prior methods that use boundary points or dedicated segmentation heads, ARGenSeg generates dense masks directly through visual token output and detokenization via a universal VQ‑VAE, enabling fine‑grained pixel‑level perception. The framework employs a next‑scale‑prediction strategy to parallelize token generation, resulting in faster inference while outperforming state‑of‑the‑art segmentation models on multiple datasets.
By Xiaolong Wang, Lixiang Ru, Ziyuan Huang, Kaixiang Ji, Dandan Zheng, Jingdong Chen, Jun Zhou
arXiv:2606. 08847v1 Announce Type: cross Abstract: Despite the success of image generation from text descriptions, it still faces challenges that are difficult to overcome in domains such as natural language processing (NLP) and computer vision (CV).
By Ahmed Abdelmoneim Mazrou, Haidy Maher El-Amir, Ali Hamdi
Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptual gap, we propose Visual In-context Editing, a new paradigm elevating video editing from textual instructions to multi-modal visual guidance encompassing single image, image pair, and video pair.
arXiv:2609.36680v1 Announce Type: new
Abstract: Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In v...
By Zizhao Li, Chengyi Cai, Mohammed Yaqoob Ansari, Feng Liu, Joseph West, Kourosh Khoshelham