arXiv AI
Sep 10

Emo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event Cameras

The paper introduces Emo-DVS, a large-scale, multimodal dataset combining event camera, audio, and text data for emotion recognition, designed to mitigate privacy concerns associated with RGB cameras. It proposes the Information‑Guided Gated Fusion (IGF) framework, which pre‑trains an event encoder on the dataset’s FAU subset, adaptively gates modalities to reduce noise, and aligns cross‑modal representations via mutual information maximization. Experiments show that IGF outperforms existing methods on this challenging tri‑modal benchmark.

By Jiaqi Chen, Qinfu Xu, Hao Zhuang, Liyuan Pan
arXiv Machine Learning
Sep 4

Beyond Blur: A Semantic Tri-view Pipeline for Teledermatology Gradability via Skin Micro-relief

The paper introduces the Semantic Tri-view Pipeline, an interpretable system that automatically screens teledermatology photographs for gradability by analyzing epidermal micro-relief across up to three smartphone views. It uses a lightweight DeepLabV3+ model to segment micro-relief fidelity and aggregates the resulting spatial masks with logistic regression, leveraging viewpoint redundancy to improve robustness. Evaluated on the SCIN dataset, the approach raises the AUC from 0.81 to 0.96 on optically clear cases, offering real‑time, privacy‑by‑design feedback to filter ungradable photo sets before clinician review.

By Robert Engel
arXiv Machine Learning
Sep 21

From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities

The study compares temporal deep learning models—Bidirectional LSTM, Temporal Convolutional Network, and Transformer—for physiological emotion recognition using two multimodal wearable datasets, WESAD and EmoWear. Experiments evaluate wrist-only, chest-only, and multimodal sensor configurations with participant-independent leave-one-subject-out cross-validation, and also explore ensembles, sensor ablation, sampling frequency, and saliency analysis. Results show that the best architecture varies by dataset, multimodal sensing consistently outperforms single-site configurations, and a 4 Hz sampling rate offers a cost-effective operating point.

By Desta Haileselassie Hagos, Saurav Keshari Aryal, Legand L. Burge