arXiv:2609.24215v1 Announce Type: new
Abstract: Although text-to-image models can accurately depict subjects and scenes, creators still struggle to specify the fine-grained emotions an image should c...
By Minglang Li, Yueyue Fang, Xieping Gao
arXiv:2505.24427v2 Announce Type: replace
Abstract: Accurate modeling of subjective phenomena such as emotion expression requires data annotated with authors' intentions. Commonly such data is collec...
By Christopher Bagdon, Aidan Combs, Carina Silberer, Roman Klinger
arXiv:2606. 05816v1 Announce Type: cross Abstract: T2I models cannot effectively capture sentiment from various types of text, including diaries, as they primarily focus on visual object-related patterns rather than contextual emotional understanding.
By Jihun Cho, Soo-Yeon Jeong, Sun-Young Ihm
arXiv:2602. 16161v4 Announce Type: replace-cross Abstract: Emotional expression underpins natural communication and effective human-computer interaction.
By Rong Fu, Ziming Wang, Shuo Yin, Haiyun Wei, Kun Liu, Xianda Li, Simon Fong
arXiv:2607. 14769v1 Announce Type: cross Abstract: Existing text summarization research has focused much on monologic information (e.
By Linyun Xiang, Mark Neerincx, Stephanie Tan
arXiv:2609.06188v1 Announce Type: new
Abstract: Multimodal sentiment analysis and emotion recognition in conversations demand effective modeling of heterogeneous interactions across textual, acoustic...
By Pengfei Shao, Jisheng Dang, Jiawen Fang, Ning Liu, Wencan Zhang, Bimei Wang, Jingwen Zhao, Jianhuang Lai, Qi Tian, Tat-Seng Chua
The paper introduces TIC‑Bench, a new benchmark for evaluating multimodal large language models on deeply interleaved text‑image contexts. It covers logical, temporal, and spatial association tasks, totaling 2,280 questions across eight specific types. The authors benchmarked ten state‑of‑the‑art MLLMs, finding a significant performance gap versus human experts and highlighting persistent challenges in integrating evidence across interleaved visual and textual inputs.
By Zihao Wang, Xi Xiang, Yuwen Sun, Yingyu Li, Yabo Zhang, Yihan Zeng, Fan Li, Wangmeng Zuo
MultiHuSE is a multimodal dataset featuring 2,407 high‑definition videos of 50 diverse actors delivering 1,463 text samples in four psychological humour styles—affiliative, aggressive, self‑enhancing, and self‑deprecating—plus neutral content. Each text is performed by multiple actors, allowing analysis of expressive diversity, and a subset includes emotion annotations. Baseline experiments show that multimodal fusion improves humour style classification accuracy over unimodal approaches, especially for affiliative humour.
By Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat
AffectDelta is a new image editing framework that moves beyond single emotion labels by modeling edits as transitions between eight‑dimensional emotion distributions. It uses a frozen Emotion Distribution Predictor to estimate the source state and a signed difference vector to encode the desired change, which is then translated into context‑dependent semantic and appearance modifications via a transition encoder and a diffusion backbone. The authors introduce AffectPair‑249K, a dataset of 248,841 source‑target pairs covering both cross‑category and within‑category transitions, and show that AffectDelta outperforms six baselines in affective alignment and content preservation.
By Xingzu Zhan, Lin Gu, Ruogu Fang
arXiv:2608. 16201v1 Announce Type: new Abstract: Multimodal sentiment analysis (MSA) aims to predict sentiment polarity and intensity from heterogeneous inputs such as text, audio, and vision.
By Shanshan Lin, Yuesheng Wu, Chao Chen, Yizhe Yang, Zhihao Chen, Zexian Yang, Xiangwen Liao
The paper introduces Affect-Prototype Guided Fusion (APCF), a framework for open‑vocabulary multimodal emotion recognition that handles incomplete and unsynchronized modal data. APCF builds an affect‑prototype library to model how different emotions contribute across modalities, enabling dynamic fusion of available features. The fused representations are then decoded by an LLM to generate open‑vocabulary emotion labels, achieving superior performance on OV‑MERD+ and MER‑FG datasets compared to existing methods.
By Yichi Zhang, Shenyue Wang, Jing Luo, Chunyang Yu, Xinyu Yang
arXiv:2608.20905v1 Announce Type: new
Abstract: Face-to-face audiovisual interaction is central to human communication, conveying rich emotional and social cues. However, existing multimodal dialogue...
By Yi Zheng, Yifan Xu, Yan Zhou, Hejia Chen, Chunyu Qiang, Xiaoqiang Liu, Xiaohan Li, Shenze Huang, Yue Zhang, Guoying Zhao, Pengfei Wan