arXiv:2512.23245v3 Announce Type: replace
Abstract: Recent text-to-image diffusion models have significantly improved visual quality and text alignment. However, generating a sequence of images while...
By Shin Seong Kim, Minjung Shin, Hyunin Cho, Youngjung Uh
arXiv:2502.03726v3 Announce Type: replace
Abstract: Text-to-image diffusion models are capable of generating high-quality images, but suboptimal pre-trained text representations often result in these...
By Zhenyu Zhou, Defang Chen, Can Wang, Chun Chen, Siwei Lyu
arXiv:2609.00476v1 Announce Type: new
Abstract: Graphic designers often blend visual concepts to communicate multiple ideas within a single image, leveraging positive and negative space to create bal...
By Shishi Xiao, Adam J. Coscia, David H. Laidlaw
Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing methods struggle to achieve accurate translation, faithful visual identity preservation, and easy-to-edit outputs, simultaneously.
arXiv:2609.01433v1 Announce Type: new
Abstract: Concept erasure aims to suppress unsafe, privacy-sensitive, or undesirable generations in text-to-image diffusion models while preserving benign semant...
By Qinghui Gong, Xunlei Chen, Yu-Xuan Zhang, Hua Meng, Zhengchun Zhou
arXiv:2606. 03792v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) successfully enables personalization in text-to-image generation by adapting pre-trained diffusion models to specific visual concepts and styles.
By Georgios Tsoumplekas, Stella Bounareli, Vasileios Argyriou
Graphic designers often blend visual concepts to communicate multiple ideas within a single image, leveraging positive and negative space to create balance, emphasis, and aesthetic appeal. While compu...
arXiv:2511.19811v2 Announce Type: replace-cross
Abstract: Image diversity remains a fundamental challenge for text-to-image diffusion models. Low-diversity generation often leads to repetitive output...
By Debin Meng, Chen Jin, Zheng Gao, Yanran Li, Ioannis Patras, Georgios Tzimiropoulos
HyperErase introduces a hypernetwork-based framework for concept erasure in text-to-image models, replacing static adapters with prompt-conditioned parameter synthesis. The method maps textual descriptions to LoRA updates, eliminating per-prompt gradient optimization and manual merging. A decoupled rectification strategy further stabilizes and refines the synthesized adapters, yielding improved erasure effectiveness, image quality, and semantic alignment across diverse concepts.
By Yi Sun, Xinhao Zhong, Zhiqi Zhang, Yimin Zhou, Junhao Li, Yuxia Qiao
Sidecar is a plug‑and‑play semantic augmentation module that preserves entity‑level information from a character’s initial description and injects missing semantics into later prompt embeddings, enabling training‑free character consistency in free‑form visual storytelling. It works with existing diffusion models without architectural changes and incurs negligible computational overhead. Experiments on FreeStoryBench demonstrate consistent improvements in prompt‑image alignment and character consistency across SDXL and FLUX baselines.
By Sibo Dong, Sarah Adel Bargal
AcFlow introduces an inference‑time controller for text‑to‑image diffusion transformers that transports intermediate layer activations through a learned, concept‑conditioned velocity field while keeping the base model frozen. The method allows fine‑grained style intensity control and suppression of unwanted concepts, achieving superior style–content trade‑offs compared to baselines and generalizing to unseen concepts without per‑concept fitting. Experiments demonstrate improved style alignment and qualitative suppression of diverse concepts where direct prompting fails.
By Junran Wang, Zehao Jin, Tianyu Luan, Xinjie Shen
arXiv:2603. 28762v2 Announce Type: replace-cross Abstract: Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of variety, converging on a narrow set of visual solutions for any given prompt.
By Omer Dahary, Benaya Koren, Daniel Garibi, Daniel Cohen-Or