arXiv:2609.24215v1 Announce Type: new
Abstract: Although text-to-image models can accurately depict subjects and scenes, creators still struggle to specify the fine-grained emotions an image should c...
By Minglang Li, Yueyue Fang, Xieping Gao
arXiv:2606. 13247v1 Announce Type: new Abstract: Text-to-image diffusion models have achieved impressive results in synthesizing high-quality images from natural language prompts.
By Emna Othmen, Mohamed Yassine Landolsi, Lotfi Ben Romdhane
arXiv:2607. 10165v1 Announce Type: cross Abstract: Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the target emotion.
By Dexiang Hong, Yijie Guo, Weidong Chen, Xinyan Liu, Zixuan Zou, Zhendong Mao, Yongdong Zhang
The paper introduces the Mult2EMo dataset, which gathers annotations from both authors and readers on multimodal social media posts and the real‑world events that triggered them. It investigates how well readers can reconstruct the authors’ emotional experience from the post content, emphasizing the importance of both text and image modalities. The study finds that accurate emotion reconstruction is possible but remains challenging, especially when images dominate the expression and when understanding the triggering event is essential.
By Christopher Bagdon, Carina Silberer, Roman Klinger
AffectDelta is a new image editing framework that moves beyond single emotion labels by modeling edits as transitions between eight‑dimensional emotion distributions. It uses a frozen Emotion Distribution Predictor to estimate the source state and a signed difference vector to encode the desired change, which is then translated into context‑dependent semantic and appearance modifications via a transition encoder and a diffusion backbone. The authors introduce AffectPair‑249K, a dataset of 248,841 source‑target pairs covering both cross‑category and within‑category transitions, and show that AffectDelta outperforms six baselines in affective alignment and content preservation.
By Xingzu Zhan, Lin Gu, Ruogu Fang
AffectDelta is a new image editing framework that moves beyond single emotion labels by modeling edits as transitions between eight‑dimensional emotion distributions. It uses a frozen Emotion Distribution Predictor to estimate the source image’s affective state and encodes the signed difference to guide a diffusion backbone that applies context‑dependent semantic and appearance changes. The authors created a large AffectPair‑249K dataset of source‑target pairs and show that AffectDelta outperforms six baselines in both affective alignment and content preservation, with ablation studies supporting their design choices.
arXiv:2608. 11452v1 Announce Type: cross Abstract: Text-to-image (T2I) models are increasingly asked to illustrate literary and cultural content, yet we cannot measure how well an image renders the meaning of a poem.
By Haoqi Hu, Tongji Luo, Li Zhang, Boning Zhou
Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personalized image generation, where models must align outputs with a user's implicit visual preferences based on a few historically preferred images and a short prompt.
arXiv:2510.12041v3 Announce Type: replace
Abstract: Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models often struggle with simple or underspecifie...
By Ruibo Chen, Jiacheng Pan, Heng Huang, Zhenheng Yang
arXiv:2608.22329v1 Announce Type: cross
Abstract: Emotion-aware artistic image generation requires a model to satisfy semantic content, artistic style, and target emotion simultaneously. The key chal...
By Qianqian Tang, Jiayi Gao, Ting Lei, Yang Liu
CompArt introduces a new approach to aesthetic alignment in text-to-image generation by using the Principles of Art (PoA) such as Balance, Rhythm, and Emphasis to define explicit compositional constraints. The authors create a large dataset of 80,032 WikiArt images, each annotated with PoA analyses generated by a multimodal LLM, and present ArtDapter, a lightweight adapter that steers a pretrained diffusion model along ten PoA dimensions while preserving semantic fidelity. Experiments demonstrate that CompArt outperforms strong baselines in adhering to PoA controls under a dual evaluation protocol.
By Zhe Jin, Tat-Seng Chua
ChatGPT Images 2. 0 introduces a state-of-the-art image generation model with improved text rendering, multilingual support, and advanced visual reasoning.