Personalizing text-to-image diffusion models to render several specific subjects in a coherent image remains challenging: the model must preserve each subject's identity while keeping the scene spatially and visually coherent. Methods that fuse independently trained concept adapters in a shared weight space (via federated averaging, gradient fusion, or orthogonality constraints) suffer from identity confusion and style bleeding and require joint retraining.
arXiv:2510.23515v3 Announce Type: replace
Abstract: This paper proposes FreeFuse, a training-free framework for multi-subject text-to-image generation through automatic fusion of multiple subject LoR...
By Yaoli Liu, Yao-Xiang Ding, Kun Zhou
arXiv:2609.01433v1 Announce Type: new
Abstract: Concept erasure aims to suppress unsafe, privacy-sensitive, or undesirable generations in text-to-image diffusion models while preserving benign semant...
By Qinghui Gong, Xunlei Chen, Yu-Xuan Zhang, Hua Meng, Zhengchun Zhou
arXiv:2608. 14172v1 Announce Type: cross Abstract: Text-to-image diffusion models have two major drawbacks that severely limit their practical utility: (1) standard models lack an intrinsic mechanism for continuous, concept-specific guidance (e.
By Nikolai R\"ohrich, Isabell Hans, Felix Krause, Bj\"orn Ommer
arXiv:2606. 03792v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) successfully enables personalization in text-to-image generation by adapting pre-trained diffusion models to specific visual concepts and styles.
By Georgios Tsoumplekas, Stella Bounareli, Vasileios Argyriou
arXiv:2605. 00273v2 Announce Type: replace-cross Abstract: Text-to-image diffusion models achieve impressive visual fidelity, yet they remain unreliable in multi-object generation.
By Yujin Jeong, Arnas Uselis, Iro Laina, Seong Joon Oh, Anna Rohrbach
arXiv:2507.00754v3 Announce Type: replace
Abstract: The integration of Large Language Model (LLMs) blocks with Vision Transformers (ViTs) holds immense promise for vision-only tasks by leveraging the...
By Selim Kuzucu, Muhammad Ferjad Naeem, Anna Kukleva, Federico Tombari, Bernt Schiele
arXiv:2608.29233v1 Announce Type: new
Abstract: We present the winning solution to the ACM Multimedia 2026 Grand Challenge on Single-Image Guided Multi-Angle Image Synthesis. It ranks first among 293...
By Jie Li, Xingchen Zou, Yuxuan Liang
arXiv:2606. 13288v1 Announce Type: cross Abstract: Contrastively trained vision-language models like CLIP, have made remarkable progress in learning joint image-text representations, but still face challenges in compositional understanding.
By Wei Li, Zhen Huang, Xinmei Tian
arXiv:2608. 03135v1 Announce Type: cross Abstract: Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts.
By Ning Zhu, An Chen, Mengfei Zhao, Juntao Xu, Jingze Liang, Boyuan Gu, Liang-Jian Deng
arXiv:2601. 03100v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) typically rely on a single late-layer feature from a frozen vision encoder, leaving the encoder's rich hierarchy of visual cues under-utilized.
By Chenchen Lin, Sanbao Su, Rachel Luo, Yuxiao Chen, Yan Wang, Marco Pavone, Fei Miao
arXiv:2606.13971v3 Announce Type: replace
Abstract: While personalizing Image-to-Video (I2V) diffusion models with specific visual effects is increasingly demanded for high-end generation, current pr...
By Xiaomeng Yang, Yanyu Li, Gordon Guocheng Qian, Ivan Skorokhodov, Viacheslav Ivanov, Avalon Vinella, Xuan Zhang, Yanzhi Wang, Sergey Tulyakov, Anil Kag