arXiv:2609.37198v1 Announce Type: new
Abstract: Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few r...
By Haoran He, Runyuan Cai, Yiming Wang, Lin Yu, Xiaodong Zeng
arXiv:2609.01433v1 Announce Type: new
Abstract: Concept erasure aims to suppress unsafe, privacy-sensitive, or undesirable generations in text-to-image diffusion models while preserving benign semant...
By Qinghui Gong, Xunlei Chen, Yu-Xuan Zhang, Hua Meng, Zhengchun Zhou
HyperErase introduces a hypernetwork-based framework for concept erasure in text-to-image models, replacing static adapters with prompt-conditioned parameter synthesis. The method maps textual descriptions to LoRA updates, eliminating per-prompt gradient optimization and manual merging. A decoupled rectification strategy further stabilizes and refines the synthesized adapters, yielding improved erasure effectiveness, image quality, and semantic alignment across diverse concepts.
By Yi Sun, Xinhao Zhong, Zhiqi Zhang, Yimin Zhou, Junhao Li, Yuxia Qiao
CompArt introduces a new approach to aesthetic alignment in text-to-image generation by using the Principles of Art (PoA) such as Balance, Rhythm, and Emphasis to define explicit compositional constraints. The authors create a large dataset of 80,032 WikiArt images, each annotated with PoA analyses generated by a multimodal LLM, and present ArtDapter, a lightweight adapter that steers a pretrained diffusion model along ten PoA dimensions while preserving semantic fidelity. Experiments demonstrate that CompArt outperforms strong baselines in adhering to PoA controls under a dual evaluation protocol.
By Zhe Jin, Tat-Seng Chua
arXiv:2506.02015v4 Announce Type: replace
Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still strug...
By Yoonjin Oh, Yongjin Kim, Hyomin Kim, Donghwan Chi, Sungwoong Kim
arXiv:2511.22245v2 Announce Type: replace
Abstract: Personalizing text-to-image diffusion models extends pretrained models to represent novel user-specific concepts from only a few reference images....
By Seoyun Yang, Gihoon Kim, Taesup Kim
arXiv:2510.12041v3 Announce Type: replace
Abstract: Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models often struggle with simple or underspecifie...
By Ruibo Chen, Jiacheng Pan, Heng Huang, Zhenheng Yang
arXiv:2605.19750v2 Announce Type: replace
Abstract: Visual autoregressive (VAR) models have recently emerged as an efficient paradigm for text-to-image generation, yet their personalization capabilit...
By Junhao Li, Xinhao Zhong, Yi sun, Yuxia Qiao, Bin Chen, Shu-Tao Xia, Yaowei Wang
Multimodal Language Models as Text-to-Image Model Evaluators presents MT2IE, a framework where a multimodal large language model generates evaluation prompts and scores images, achieving higher correlation with human judgment than prior metrics. MT2IE recovers official T2I model rankings using only 20 prompts—far fewer than traditional benchmarks—and adapts prompts to each model’s performance, maintaining informative scoring ranges. The approach demonstrates that dynamic, interactive evaluation can replace static benchmarks as T2I models improve.
By Jiahui Chen, Candace Ross, Reyhane Askari-Hemmat, Koustuv Sinha, Melissa Hall, Amy Zhang, Michal Drozdzal, Adriana Romero-Soriano
arXiv:2606. 26668v1 Announce Type: cross Abstract: Video customization based on Text-to-Video (T2V) models aims to learn specific features from reference data to generate controllable videos.
By Xuancheng Xu, Gengyun Jia, Bing-Kun Bao
The paper introduces Redemption Score (RS), a multi‑modal evaluation framework for image captioning that combines three complementary signals: Mutual Information Divergence for global image‑text alignment, DINO‑based perceptual similarity of cycle‑generated images for visual grounding, and LLM text embeddings for contextual similarity to human references. RS fuses these signals to provide a more holistic assessment, achieving a Kendall‑τ of 58.42 on Flickr8k and outperforming most prior methods. The framework demonstrates consistent performance across Conceptual Captions and MS COCO, offering a robust evaluation that captures both visual accuracy and text quality.
By Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi
arXiv:2606. 08841v1 Announce Type: new Abstract: Text-to-image diffusion models are increasingly deployed in open-ended creative contexts, yet their outputs remain impersonal, optimized for aggregate aesthetics rather than individual taste.
By Harini SI, Somesh Singh, Yaman Kumar Singla, David Doermann, Rajiv Ratn Shah