arXiv Machine Learning

To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection

arXiv:2606. 05931v1 Announce Type: cross Abstract: When retrieving a person from a video archive by voice and face, should the system be multimodal or not?

arXiv Computer Vision
Sep 15

Query-Conditioned Spherical Centroid Aggregation for Multimodal Retrieval

The paper introduces SCALAR, a query‑conditioned spherical centroid aggregator that assigns relevance‑based weights to each available modality before computing a spherical centroid. SCALAR supports arbitrary modality subsets, is trained with rank‑8 LoRA adapters on masked views, and achieves positive aggregation gains on four of five benchmarks, outperforming prior symmetric aggregators. With only 4.8 million trainable parameters, SCALAR attains the highest text‑to‑video R@1 on three benchmarks and surpasses the released GRAM checkpoint under test‑time modality dropout by 3.2 to 10.9 R@1.

By Ambuj Mehrish, Anindya Nag, Sebastiano Vascon
arXiv AI
Jul 10

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

arXiv:2607. 07907v1 Announce Type: cross Abstract: With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training data.

By Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu, Vaidehi Patil
Hugging Face Trending Papers
Jul 23

Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features

Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting. We hypothesize that audiovisual expression dynamics carry identity-discriminative information complementary to static appearance, and that extracting this signal requires multimodal representations robust to the variable input quality of in-the-wild video.

arXiv Machine Learning
Sep 15

ReH-FUSE: Reliability-Aware Hierarchical Fusion of Experts for Multimodal Emotion Recognition in Conversation

ReH-FUSE is a reliability‑aware hierarchical fusion framework for multimodal emotion recognition in conversation. It uses a decision‑level router to first compare the relative preference between text and audio, then balances this unimodal mixture with a cross‑modal expert, thereby separating unimodal competition from cross‑modal selection. Experiments on IEMOCAP and MELD show that ReH-FUSE achieves state‑of‑the‑art weighted and macro F1 scores, and ablation studies confirm that learned routing outperforms uniform expert averaging and benefits from cross‑modal interaction.

By Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen
arXiv AI
Sep 10

RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavioral Monitoring in Surveillance Systems

The paper introduces RAFM-SER++, a lightweight multimodal speech emotion recognition framework designed for real‑time surveillance systems. It replaces heavy bidirectional cross‑modal transformers with an asymmetric Residual Attention Fusion Mechanism that injects affective speech cues into text representations via a one‑directional residual attention pathway. Experiments on IEMOCAP and ESD show RAFM‑SER++ outperforms the HuBERT‑Base baseline and MemoCMT, reducing trainable parameters by over 60%, achieving 79.60 it/s inference speed, and reaching BACC scores of 81.10% on IEMOCAP and 95.39% on ESD.

By Ngo Truong Dinh, Tung-Lam Bui, Chi-Trung Duong, Vien Nguyen Thi, Viet-Anh Nguyen, Phuc-Lu Le