arXiv:2609.24215v1 Announce Type: new
Abstract: Although text-to-image models can accurately depict subjects and scenes, creators still struggle to specify the fine-grained emotions an image should c...
By Minglang Li, Yueyue Fang, Xieping Gao
AffectDelta is a new image editing framework that moves beyond single emotion labels by modeling edits as transitions between eight‑dimensional emotion distributions. It uses a frozen Emotion Distribution Predictor to estimate the source state and a signed difference vector to encode the desired change, which is then translated into context‑dependent semantic and appearance modifications via a transition encoder and a diffusion backbone. The authors introduce AffectPair‑249K, a dataset of 248,841 source‑target pairs covering both cross‑category and within‑category transitions, and show that AffectDelta outperforms six baselines in affective alignment and content preservation.
By Xingzu Zhan, Lin Gu, Ruogu Fang
arXiv:2608.22329v1 Announce Type: cross
Abstract: Emotion-aware artistic image generation requires a model to satisfy semantic content, artistic style, and target emotion simultaneously. The key chal...
By Qianqian Tang, Jiayi Gao, Ting Lei, Yang Liu
AffectDelta is a new image editing framework that moves beyond single emotion labels by modeling edits as transitions between eight‑dimensional emotion distributions. It uses a frozen Emotion Distribution Predictor to estimate the source image’s affective state and encodes the signed difference to guide a diffusion backbone that applies context‑dependent semantic and appearance changes. The authors created a large AffectPair‑249K dataset of source‑target pairs and show that AffectDelta outperforms six baselines in both affective alignment and content preservation, with ablation studies supporting their design choices.
arXiv:2606. 13247v1 Announce Type: new Abstract: Text-to-image diffusion models have achieved impressive results in synthesizing high-quality images from natural language prompts.
By Emna Othmen, Mohamed Yassine Landolsi, Lotfi Ben Romdhane
arXiv:2609.12830v1 Announce Type: new
Abstract: Continuous emotion control in text-to-image generation requires a model to improve affective alignment without changing the objects, layout, or scene d...
By Jisheng Dang, Zhenxuan Wang, Bin Li, Ronghao Lin, Bin Hu, Tat-Seng Chua
arXiv:2607. 10678v1 Announce Type: new Abstract: Emotional intelligence enables humans to recognize emotions, infer their causes, reason about interventions, and modify their environment to achieve desired affective states.
By Qing Lin, Mengmi Zhang
The paper introduces DAN, a training‑free inference‑time framework that improves affective reasoning in multimodal large language models. It combines a Hierarchical Emotional Reasoning Chain (HERC) to better capture fine‑grained visual cues and a Contrastive Discriminative Visual Pruning (CDVP) module to isolate discriminative tokens for semantically similar emotions. Experiments show significant gains, notably a +10.47% improvement on the WebEmo25 benchmark with Qwen3‑VL‑8B‑Instruct.
By Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu, Zhendong Mao
arXiv:2609.38182v1 Announce Type: cross
Abstract: Avatar-based multimodal empathetic response generation has emerged as a pivotal capability in human-centric systems, aiming to recognize user emotion...
By Xiaolin Chen, Xuemeng Song, Jinlan Fu, Weili Guan, Mong-Li Lee, Wynne Hsu
Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases.
arXiv:2606. 05816v1 Announce Type: cross Abstract: T2I models cannot effectively capture sentiment from various types of text, including diaries, as they primarily focus on visual object-related patterns rather than contextual emotional understanding.
By Jihun Cho, Soo-Yeon Jeong, Sun-Young Ihm
SenseNova-U1.5 is an 8B‑MoT native unified multimodal model that can understand, reason about, and generate visual content without using an encoder or VAE. It improves visual fidelity and text rendering through spatially coherent patch reconstruction, large‑scale training on curated generation and editing data, and native resolutions up to 4K. Post‑training, specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing are optimized and distilled into a multi‑expert framework, yielding advances in image fidelity, complex composition, multi‑reference editing, and instruction following.
By Haiwen Diao, Jiahao Wang, Chenjing Ding, Hanming Deng, Jiangnan Chen, Ruixi Zhang, Ruohui Wang, Wenwen Tong, Xiangyu Fan, Yubo Wang, Yue Zhu, Yuwei Niu, Zhengqi Bai, Zhiqian Lin, Zhitao Yang, Zhongang Cai, Bo Yang, Chen Feng, Chengguang Lv, Guangjia Liu, Guanlin Wang, Hanyu Zhang, Haojia Yu, Hongcan Xiao, Hongli Wang, Huan Wu, Huaping Zhong, Jian Fang, Jianan Fan, Jiaqi Li, Jiefan Lu, Jing Zuo, Jingcheng Ni, Junxiang Xu, Linjun Dai, Mutian Xu, Peishen Yan, Penghao Wu, Ruijie Mao, Ruisi Wang, Shihao Bai, Shuang Yang, Shuya Yang, Shuyan Zheng, Silei Wu, Siying Li, Tao Chu, Tianbo Zhong, Tongxi Zhou, Weichao Luo, Weichen Fan, Wenhao Jia, Wenjie Gao, Xiangli Kong, Yan Li, Yang Yong, Zimo Wen, Zixuan Qian, Wenxiu Sun, Ruihao Gong, Quan Wang, Lewei Lu, Lei Yang, Ziwei Liu, Dahua Lin