Multimodal sarcasm detection aims to identify sarcastic intent from multimodal content, where inconsistencies between literal meaning and contextual cues often signal irony. This task has attracted increasing research attention.
arXiv:2607. 16076v1 Announce Type: cross Abstract: Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from incongruity between textual and visual information rather than from either modality alone.
By Bhavana Verma, Priyanka Meel, Dinesh Kumar Vishwakarma
arXiv:2607. 06611v1 Announce Type: cross Abstract: Automatically recognizing the sentiment, positive or negative, from speech is a challenging task, requiring both the analysis of vocal inflections and the interpretation of uttered words.
By Andrei-George Durdun, Victor Constantinescu, Radu Tudor Ionescu
arXiv:2607. 19011v1 Announce Type: cross Abstract: Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description.
By Tuo Liang, Zhe Hu, Disheng Liu, Jing Li, Yu Yin
arXiv:2606. 15694v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in understanding complex multimodal content.
By Hangling Xie
arXiv:2608. 19971v1 Announce Type: new Abstract: Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues.
By Zhifa Geng, Subin Huang, Hao Guo, Junjie Chen, Sanmin Liu, Chao Kong
arXiv:2608. 20019v1 Announce Type: new Abstract: Incomplete multimodal sentiment analysis has garnered significant attention in recent years.
By Kaixin Xu, NaiJin Liu, Yulin Kang, Tangyue Jin, Zixuan Yu, Wenxi Zhao, Yibei Liu, Qianle Zhang, Yangyang Wu, Mengying Zhu, Meng Xi
arXiv:2608. 04054v1 Announce Type: cross Abstract: Multimodal intent recognition requires understanding not only what textual, acoustic, and visual signals share, but also how they disagree.
By Mohnish Raj, Suraj Kumar, Soumi Chattopadhayay, Chandranath Adak, Ayan Dutta
arXiv:2607. 15442v1 Announce Type: new Abstract: Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where humor, sarcasm, and harmful intent coexist.
By Shanhong Liu, Pai Chet Ng, De Wen Soh, Malika Meghjani, Konstantinos N. Plataniotis
arXiv:2601. 07565v2 Announce Type: replace-cross Abstract: Multimodal emotion understanding requires the integration of heterogeneous data sources, including text, audio, and visual modalities, while simultaneously addressing discrete emotion recognition and continuous sentiment analysis.
By Jiaqi Qiao, Xinran Li, Yifan Lyu, Xiujuan Xu, Liu Yu
arXiv:2606. 02679v1 Announce Type: new Abstract: Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed.
By Jiyuan Liu, Liangwei Nathan Zheng, Wei Emma Zhang, Xinpei Wang, Weitong Chen
arXiv:2509. 25773v3 Announce Type: replace-cross Abstract: AI models capable of comprehending humor hold real-world promise -- for example, enhancing engagement in human-machine interactions.
By Zhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang, Qinrong Cui, Wei Bi, Songchun Zhu, Bo Zhao, Zilong Zheng