arXiv:2606. 09653v1 Announce Type: new Abstract: Learned representations across models and modalities often exhibit striking structural similarities, suggesting shared underlying concept decompositions.
By Gr\'egoire Dhimo\"ila, Victor Boutin, Agustin Martin Picard, Thomas Fel, Thomas Serre
arXiv:2603. 26798v2 Announce Type: replace-cross Abstract: Vision-language model (VLM) encoders such as CLIP enable strong retrieval and zero-shot classification in a shared image-text embedding space, yet the semantic organization of this space is rarely inspected.
By Gesina Schwalbe, Mert Keser, Moritz Bayerkuhnlein, Edgar Heinert, Annika M\"utze, Marvin Keller, Sparsh Tiwari, Georgii Mikriukov, Diedrich Wolter, Jae Hee Lee, Matthias Rottmann
arXiv:2607. 00402v1 Announce Type: cross Abstract: Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts.
By Adeel Yousaf, Soumik Ghosh, James Beetham, Amrit Singh Bedi, Mubarak Shah
arXiv:2607. 06432v1 Announce Type: cross Abstract: Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and safety regulations, deployed systems must be able to suppress unwanted concepts after training.
By Naveen George, Naoki Murata, Yuhta Takida, Konda Reddy Mopuri, Yuki Mitsufuji
arXiv:2609.36348v1 Announce Type: cross
Abstract: Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas th...
By Xiaoyu Wu, Yifei Wang, Chen Wei
arXiv:2607. 08337v1 Announce Type: new Abstract: Diffusion unlearning is essential for mitigating the generation of harmful or copyrighted content in text-to-image models.
By Siyuan Wen, Jiahao Zeng, Ningning Ding
Diffusion unlearning is essential for mitigating the generation of harmful or copyrighted content in text-to-image models. Current diffusion unlearning techniques determine the model update direction by either using alternatives of the target concept as an anchor or using empty prompts.
arXiv:2503. 07265v4 Announce Type: replace-cross Abstract: Text-to-Image (T2I) models are capable of generating high-quality artistic creations and visual content.
By Yuwei Niu, Munan Ning, Mengren Zheng, Weiyang Jin, Bin Lin, Peng Jin, Jiaqi Liao, Chaoran Feng, Fanqing Meng, Kunpeng Ning, Bin Zhu, Li Yuan
arXiv:2606. 15819v1 Announce Type: cross Abstract: The rapid progress of visual autoregressive (VAR) models has unlocked a transformative frontier for high-fidelity text-to-image synthesis, while heightening concerns over the safety alignment of generated content.
By Siya Yang, Nanxiang Jiang, Zhaoxin Fan, Yunfeng Diao
arXiv:2511.22245v2 Announce Type: replace
Abstract: Personalizing text-to-image diffusion models extends pretrained models to represent novel user-specific concepts from only a few reference images....
By Seoyun Yang, Gihoon Kim, Taesup Kim
arXiv:2606. 10892v1 Announce Type: cross Abstract: To showcase products, merchants often incur substantial costs creating high-quality display images.
By Yihao Zhao, Xuan Han, Bin He, Mingyu You
The paper introduces Mapping the Concept Landscape (MCL), a framework that replaces high‑dimensional feature embeddings with explicit sample‑level graphs of entities, events, and attributes for image‑caption pairs. By aggregating these graphs into a dataset‑level graph, MCL captures the global distribution of semantic concepts and identifies rare concepts. A greedy algorithm then selects samples to maximize coverage of under‑represented concepts, achieving better pruning efficiency and providing a transparent audit trail.
By Dongyue Wu, Tao Ma