arXiv:2607. 14702v1 Announce Type: cross Abstract: Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, acoustic, facial, and contextual patterns, while top-performing systems often rely on computationally expensive ensembles.
By Elena Ryumina (St. Petersburg Federal Research Center of the Russian Academy of Sciences), Maxim Markitantov (St. Petersburg Federal Research Center of the Russian Academy of Sciences), Alexandr Axyonov (St. Petersburg Federal Research Center of the Russian Academy of Sciences), Fedor Shchetinin (HSE University, St. Petersburg, Russia), Timur Abdulkadirov (St. Petersburg Federal Research Center of the Russian Academy of Sciences), Dmitry Ryumin (St. Petersburg Federal Research Center of the Russian Academy of Sciences), Alexey Karpov (St. Petersburg Federal Research Center of the Russian Academy of Sciences)
arXiv:2607.13345v2 Announce Type: replace
Abstract: We present a frame-independent audio-text system for the 3rd Ambivalence/Hesitancy Video Recognition Challenge at the 11th Affective & Behavior Ana...
By Luiz F. B. F. Martins, Rodrigo W. Pisaia, Matheus M. Girardi, Isabella V. Berkembrock, Jo\~ao A. Almeida, Andre G. Hochuli, Rayson Laroca, Alceu S. Britto Jr
The paper introduces RAFM-SER++, a lightweight multimodal speech emotion recognition framework designed for real‑time surveillance systems. It replaces heavy bidirectional cross‑modal transformers with an asymmetric Residual Attention Fusion Mechanism that injects affective speech cues into text representations via a one‑directional residual attention pathway. Experiments on IEMOCAP and ESD show RAFM‑SER++ outperforms the HuBERT‑Base baseline and MemoCMT, reducing trainable parameters by over 60%, achieving 79.60 it/s inference speed, and reaching BACC scores of 81.10% on IEMOCAP and 95.39% on ESD.
By Ngo Truong Dinh, Tung-Lam Bui, Chi-Trung Duong, Vien Nguyen Thi, Viet-Anh Nguyen, Phuc-Lu Le
We present an audio-text system for the Ambivalence/Hesitancy Video Recognition Challenge of the 11th ABAW Competition. The method excludes visual frames and represents each video as overlapping 5-second windows aligned with transcript timestamps.
The paper introduces the Modality Discrepancy Transformer (MDT), a model designed to detect ambivalence and hesitancy in clinical videos by capturing cross‑modal disagreement across facial, vocal, and linguistic signals. MDT expands a 6‑token representation to 9 tokens that include modality embeddings, absolute‑difference features, and Hadamard‑product discrepancy features, which are processed through Transformer self‑attention with FiLM‑based text conditioning and LoRA fine‑tuning. On the BAH dataset from the 3rd ABAW Challenge, MDT achieves a Macro F1 score of 0.7408 on the labelled test split and 0.7368 on the private leaderboard, surpassing the strongest baseline by over 10 points while training in under 20 minutes on a single GPU.
By Shiyu Luo, Yu Wang, Jiawen Huang, Zhaoxiang Xiao, Chenxi Huang, Qi Zhang, Bin Liu
The paper introduces a multimodal emotion recognition framework that combines audio and visual feature extraction with an attention-based fusion strategy. Audio features include Wav2Vec2 embeddings, MFCCs, and statistical acoustic descriptors, fused via a BiLSTM, while video features are extracted using a ResNet50-BiLSTM architecture. A multi-head attention mechanism fuses these modalities, and experiments on MELD and IEMOCAP show significant accuracy and robustness gains, especially in unbalanced data settings.
By Xu Lin, Ke Wang, Hui Kang, Xinying Wang