arXiv:2608.21229v1 Announce Type: new
Abstract: Omnimodal generation is central to a wide range of content creation and editing applications. In-context conditioning is essential to this paradigm. It...
By Yangshuai Liu, Zheming Li, Jiaao Li, Kang He, Ziliang Lai, Zhitai Liu, Chengru Song
arXiv:2607.18227v2 Announce Type: replace
Abstract: In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and imag...
By Dingyun Zhang, Lixue Gong, Wei Liu
arXiv:2609.24510v2 Announce Type: replace
Abstract: Recent advances in vision foundation models (VFMs) have shown remarkable capabilities across diverse unimodal visual tasks. However, adapting VFMs...
By Xiaoqiang Lu, Licheng Jiao, Lingling Li, Yuting Yang, Long Sun, Wenping Ma, Xu Liu, Fang Liu
arXiv:2609.37198v1 Announce Type: new
Abstract: Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few r...
By Haoran He, Runyuan Cai, Yiming Wang, Lin Yu, Xiaodong Zeng
RefDiT is a new framework for reference-guided image generation that addresses the shortcomings of previous methods when handling complex scenes with multiple objects. It introduces local region guidance by decomposing a single identifier token into attribute-level signals, allowing the model to learn correspondences between tokens and specific regions of a reference image. The approach incorporates a low-rank adapter (LoRA) within a diffusion transformer (DiT) to adjust the inference prompt based on user-provided guidance context, thereby enabling more precise local attribute control.
By Rameshwar Mishra, Srikrishna Karanam, A V Subramanyam
Video Diffusion Transformers (DiTs) spend most of their compute inside the Self-Attention operation, whose cost grows quadratically, $\mathcal{O}(n^2)$, with the number of latent tokens $n$. For the task of video generation, the token count is large, so this term dominates runtime and memory, and thereby caps the resolution and duration we can generate.