SafeStyle is a training‑free framework that injects calibrated style residuals into frozen diffusion models for reference‑guided stylization. It estimates style‑supported and content‑associated subspaces from small calibration sets, then transports purified style evidence across adaptive spatial granularity while limiting its influence with a residual‑norm budget. Experiments on texture‑ and geometry‑dominant styles show high DINO style similarity (0.432–0.474) with minimal semantic leakage (0.8%).
By Zhangping Yang, Min Li, Song Yan, Rong Gao, Xinliang Bi, Guanye Xiong, Yujie He
arXiv:2608. 19719v1 Announce Type: cross Abstract: Reference-based diffusion stylization requires separating target geometry from transferable appearance.
By Jingtao Zhang, Haorui Gao, Youqing Liang, Zeming Liu
arXiv:2602.08059v2 Announce Type: replace-cross
Abstract: Text-to-image diffusion models can reproduce specific artists visual styles at extremely low cost, raising copyright and deployment safety co...
By Tong Zhang, Ru Zhang, Jianyi Liu
Reference-based diffusion stylization requires separating target geometry from transferable appearance. Existing tuning-based methods often rely on aligned content-style-target triplets or auxiliary visual encoders, which increases data cost and can transfer unintended scene structure from the style reference.
arXiv:2610.02044v1 Announce Type: new
Abstract: Recent advances in rectified flow-based image-to-3D generative models have enabled high-fidelity 3D asset generation. Building on this, a growing line...
By Tao Wu, Alexandra Gomez-Villa, Senmao Li, Yaxing Wang, Joost van de Weijer, Kai Wang
Abstract‑LoRA introduces a lightweight LoRA training approach that targets specific U‑Net blocks in diffusion models to improve single‑image style transfer. By refining block selection, adding more blocks, and using clustering‑based style abstraction, it better disentangles and balances style and content compared to prior methods like B‑LoRA. Experiments show that the method produces more harmonious artistic images while quantitatively preserving both style and content.
By Xinglin Hu
GraftSR is a diffusion-based super‑resolution framework that uses reference images of the same object to guide texture restoration, mitigating hallucination. It introduces a dual‑mask reference guidance mechanism to decouple texture extraction from application, avoiding reliance on spatial alignment. The authors also release TexRefSR‑141K, a large dataset of reference pairs with spatial masks, and show that GraftSR outperforms existing methods on the TexRefSR‑Eval benchmark, reducing LPIPS by 20.2%.
By Qifan Yu, Haoran Bai, Zongyao He, Weijie He, Sibin Deng, Honggang Qi, Ying Chen
MAST (Mask‑Guided Attention Control for Training‑Free Regional‑Multi Style Transfer) is a framework that enables diffusion models to apply multiple reference styles to user‑specified regions of a content image without any training or optimization. It introduces logit‑level attention mass allocation, sharpness‑aware temperature scaling, and discrepancy‑aware detail injection to address mass allocation, selectivity, and detail loss problems in regional‑multi style transfer. Experiments with two to five styles show that MAST outperforms baselines in ArtFID, FID, and R‑FID, achieving high regional style fidelity, content preservation, and scalability.
By Dongkyung Kang, Jaeyeon Hwang, Junseo Park, Minji Kang, Yeryeong Lee, Beomseok Ko, Hanyoung Roh, Jeongmin Shin, Hyeryung Jang
Artistic image synthesis aims to recreate the expressive visual identity of a target artist, yet existing methods often fail to capture an artist's global style. Conventional style transfer methods transfer the style of one or a few reference artworks to a content image in a One-to-One manner, making them effective for artwork-level stylization but limited in representing the broader stylistic distribution of an artist.
arXiv:2610.01969v1 Announce Type: new
Abstract: Concept erasure aims to remove a target concept, such as a copyrighted style, a recognizable character, or unsafe content, from a pretrained text-to-im...
By Yongliang Wu, Haori Lu, Yulun Wu, Jinqi Luo, Xingyu Zhu, Yaoyao Liu
ReGain is a training‑free correction that improves subject fidelity in text‑to‑image diffusion models personalized with synthetic images. The authors show that fine‑tuning on synthetic images degrades fidelity due to inflated classifier‑free guidance, especially at high frequencies. ReGain measures this inflation per frequency band and scales it down during sampling, closing 51‑64% of the fidelity gap on Stable Diffusion v1.5 and improving performance on SDXL and SD 3.5 while preserving text alignment.
By Shubhang Bhatnagar, Ishan Bhatnagar, Viraj Shah, Narendra Ahuja
Despite remarkable progress in text-guided image editing, generative models frequently fail to preserve visual object consistency, defined as the preservation of a subject's key attributes throughout the editing process. We address this limitation through three contributions.