arXiv Computer Vision

An Intelligent Decision Support System for Emotion Monitoring using Microscopic Fixational Dynamics

arXiv AI
Sep 10

Emo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event Cameras

The paper introduces Emo-DVS, a large-scale, multimodal dataset combining event camera, audio, and text data for emotion recognition, designed to mitigate privacy concerns associated with RGB cameras. It proposes the Information‑Guided Gated Fusion (IGF) framework, which pre‑trains an event encoder on the dataset’s FAU subset, adaptively gates modalities to reduce noise, and aligns cross‑modal representations via mutual information maximization. Experiments show that IGF outperforms existing methods on this challenging tri‑modal benchmark.

By Jiaqi Chen, Qinfu Xu, Hao Zhuang, Liyuan Pan
arXiv Machine Learning
Sep 4

Beyond Blur: A Semantic Tri-view Pipeline for Teledermatology Gradability via Skin Micro-relief

The paper introduces the Semantic Tri-view Pipeline, an interpretable system that automatically screens teledermatology photographs for gradability by analyzing epidermal micro-relief across up to three smartphone views. It uses a lightweight DeepLabV3+ model to segment micro-relief fidelity and aggregates the resulting spatial masks with logistic regression, leveraging viewpoint redundancy to improve robustness. Evaluated on the SCIN dataset, the approach raises the AUC from 0.81 to 0.96 on optically clear cases, offering real‑time, privacy‑by‑design feedback to filter ungradable photo sets before clinician review.

By Robert Engel
arXiv Machine Learning
Sep 21

From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities

The study compares temporal deep learning models—Bidirectional LSTM, Temporal Convolutional Network, and Transformer—for physiological emotion recognition using two multimodal wearable datasets, WESAD and EmoWear. Experiments evaluate wrist-only, chest-only, and multimodal sensor configurations with participant-independent leave-one-subject-out cross-validation, and also explore ensembles, sensor ablation, sampling frequency, and saliency analysis. Results show that the best architecture varies by dataset, multimodal sensing consistently outperforms single-site configurations, and a 4 Hz sampling rate offers a cost-effective operating point.

By Desta Haileselassie Hagos, Saurav Keshari Aryal, Legand L. Burge
Hugging Face Trending Papers
Aug 19

EgoHRV: Continuous Heart Rate Variability Estimation from Egocentric Systems for Autonomic Response and Skill Assessment

EgoHRV is a method that estimates heart rate variability (HRV) and heart rate (HR) from the gaze cameras in egocentric headsets. It uses a 3D backbone and a low–high decomposition module to extract the blood volume pulse signal from gaze video, and aligns frequency‑domain representations of contact‑based and camera‑derived signals through cross‑domain pretraining. The approach achieves state‑of‑the‑art accuracy for HR and HRV estimation and, when integrated into EgoExo4D’s proficiency estimator, improves accuracy by 17.8%.

arXiv Computer Vision
Sep 18

SeetaPsych v1.0: An Open-source Computer Vision Toolkit for Behavior-based Psychological Measurement

SeetaPsych v1.0 is an open‑source computer vision toolkit that unifies and extends modules for behavior‑based psychological measurement, focusing on facial image and video analysis. It includes four core modules—emotion analysis, camera‑based heart rate estimation, screen point‑of‑gaze estimation, and scene gaze following—alongside preprocessing tools such as face, landmark, and head detection. The toolkit offers a modular Pipeline/Runner architecture, standardized Python APIs, and an interactive WebUI to enable reproducible, large‑scale research across psychology, behavioral science, and human‑computer interaction.

By Jiabei Zeng, Chiqin Li, Kaizhou Li, Fei Chang, Yong Li, Yuanhao Zhao, Dan Han, Wenqiang Yang, Xilin Chen, Shiguang Shan
arXiv AI
Aug 18

Take it Personally: The Limits of General SSL Representations for Real-Life PPG Emotion Detection

arXiv:2608. 14675v1 Announce Type: cross Abstract: While Self-Supervised Learning (SSL) effectively extracts general representations from noisy, unconstrained physiological signals such as photoplethysmography (PPG), its suitability for highly subjective tasks remains unproven.

By Dominika Kunc, Przemys{\l}aw Kazienko, Stanis{\l}aw Saganowski