arXiv AI By Mojtaba Moattari

Interpretable Multimodal Classification with Linear Discriminant Tree Ensembles

Read the original on arXiv AI →

The paper introduces Linear Discriminant Tree Ensembles (LDT, LDF, LDAB) for multimodal affect and behaviour classification, combining text, audio, and visual streams. It uses tokenization, concept clustering, and tree-based routing to balance accuracy and interpretability, and proposes a modified feature importance metric that reduces negative class bias. The ensembles outperform Multimodal Transformers and Interpretable Multimodal Routing in F1-mod and accuracy, and their feature importance aligns better with human annotations on IEMOCAP and CMU-MOSI datasets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 17

Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis

The paper proposes a Mixture-of-Bottleneck (MoB) framework for video-based multimodal sentiment analysis that treats sentiment as an ordinal regression problem, splitting it into polarity recognition and intensity prediction. MoB assigns modality‑specific latent experts to each sub‑task, learns compact, task‑relevant representations via an information bottleneck, and fuses these experts with a multimodal bottleneck routing module and hard mining strategy. Experiments on four datasets and language models demonstrate that MoB captures fine‑grained intra‑ and inter‑modal dynamics, improving performance and enabling more trustworthy localization of nuanced sentiment signals.

By Ronghao Lin, Qiaolin He, Zefeng Lu, Yichu Liu, Li Huang, Sijie Mai, Haifeng Hu, Yap-peng Tan
arXiv AI
Sep 10

RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavioral Monitoring in Surveillance Systems

The paper introduces RAFM-SER++, a lightweight multimodal speech emotion recognition framework designed for real‑time surveillance systems. It replaces heavy bidirectional cross‑modal transformers with an asymmetric Residual Attention Fusion Mechanism that injects affective speech cues into text representations via a one‑directional residual attention pathway. Experiments on IEMOCAP and ESD show RAFM‑SER++ outperforms the HuBERT‑Base baseline and MemoCMT, reducing trainable parameters by over 60%, achieving 79.60 it/s inference speed, and reaching BACC scores of 81.10% on IEMOCAP and 95.39% on ESD.

By Ngo Truong Dinh, Tung-Lam Bui, Chi-Trung Duong, Vien Nguyen Thi, Viet-Anh Nguyen, Phuc-Lu Le
arXiv AI
Sep 10

MVFA: A Multi-View Text-Guided Multimodal Fusion LLM Adapter for Sentiment Analysis and Emotion Recognition

arXiv:2609.06188v1 Announce Type: new Abstract: Multimodal sentiment analysis and emotion recognition in conversations demand effective modeling of heterogeneous interactions across textual, acoustic...

By Pengfei Shao, Jisheng Dang, Jiawen Fang, Ning Liu, Wencan Zhang, Bimei Wang, Jingwen Zhao, Jianhuang Lai, Qi Tian, Tat-Seng Chua