arXiv:2606. 14125v1 Announce Type: cross Abstract: Inversion-based image editing offers flexible and training-free control but still struggles with inversion accuracy and the trade-off between editing fidelity and background preservation.
By Zheyuan Zhan, Hongchen Li, Can Wang, Yinfei Ma, Mingzhen Huang, Ruoshi Bai, Jiawei Chen, Siwei Lyu, Defang Chen
arXiv:2603. 28762v2 Announce Type: replace-cross Abstract: Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of variety, converging on a narrow set of visual solutions for any given prompt.
By Omer Dahary, Benaya Koren, Daniel Garibi, Daniel Cohen-Or
arXiv:2507. 17853v2 Announce Type: replace-cross Abstract: Recent advances in text-to-image (T2I) generation have led to impressive visual results.
By Lifeng Chen, Jiner Wang, Zihao Pan, Beier Zhu, Xiaofeng Yang, Chi Zhang
arXiv:2602.10216v2 Announce Type: replace
Abstract: A single text prompt passed to a diffusion model yields a wide range of visual outputs determined solely by a stochastic process, leaving users wit...
By Pawe{\l} Skier\'s, Emilia Kaczmarczyk, Tomasz Trzci\'nski, Kamil Deja
arXiv:2606. 05950v1 Announce Type: new Abstract: Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models.
By Yuxiao Ye, Haoran He, Fangyuan Kong, Xintao Wang, Pengfei Wan, Kun Gai, Ling Pan
arXiv:2610.01723v1 Announce Type: new
Abstract: Text-to-image diffusion models have achieved remarkable progress in image synthesis, yet can exhibit memorization by closely reproducing individual tra...
By Hyungjun Joo, Sehwan Kim, Hyeonggeun Han, Sangwoo Hong, Jungwoo Lee
arXiv:2609.08032v1 Announce Type: cross
Abstract: We introduce FlexMoGen, a novel framework for flexible human motion synthesis conditioned on both natural language descriptions and motion style refe...
By Kai Weixian Lan, Bodie Criswell, Briana Fedkiw, Zhan Zhang, Joseph Teran, Daniel Holden
arXiv:2606. 17979v1 Announce Type: new Abstract: Existing RL post-training methods for text-to-image generation usually convert the final-image reward into a single scalar advantage and apply it with the same strength to the entire generative trajectory.
By Jinjie Shen, Wei Deng, Xian Hu, Daiguo Zhou, Jian Luan
arXiv:2606. 10892v1 Announce Type: cross Abstract: To showcase products, merchants often incur substantial costs creating high-quality display images.
By Yihao Zhao, Xuan Han, Bin He, Mingyu You
ContextAnyone is a context‑aware diffusion framework that treats a reference image as an explicitly preserved appearance anchor rather than a simple conditioning signal. By jointly reconstructing the reference image and generating the target video within a shared diffusion transformer, it provides direct supervision for maintaining identity and fine‑grained appearance throughout denoising. The method introduces asymmetric information flow and Gap‑RoPE positional representations to keep the reference stable while allowing selective access by video tokens, and demonstrates improved identity and appearance consistency on an OpenVid‑HD benchmark.
By Ziyang Mai, Yu-Wing Tai
arXiv:2609.22916v1 Announce Type: new
Abstract: Generating text-rich images from prompts requires both textual fidelity and the coherent integration of text into the surrounding image. An explicit la...
By Guanqiao Chen, Jingru Tan, Dongxing Mao, Catherine Chen, Zijian Du, Libo Qin, Hu Jian Guo, Alex Jinpeng Wang
arXiv:2607. 21529v1 Announce Type: cross Abstract: Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing.
By Yueyi Liu, Chi Zhang, Sen Cui, Miao Liu