Hugging Face Trending Papers

Rethinking the Role of Feature Engineering and Learning Strategies in Few-Shot Hidden Emotion Recognition

Read the original on Hugging Face Trending Papers →

In this paper, we present the solution developed by our team, XInsight Lab, which achieved first place in Track 3 of the 4th EI-MIGA-IJCAI Challenge with a test accuracy of 0. 76923.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computer Vision
Sep 11

Single-Stream Multi-Feature Fusion with Temporal Robustness for Gait Emotion Recognition

The paper introduces SV-GCN, a single-stream multi-feature fusion framework for 3D skeleton-based gait emotion recognition that incorporates temporal invariance. It uses intra-frame relative motion features to remove frame-rate sensitivity and embeds heterogeneous cues at shallow layers for early fusion, avoiding multi-stream complexity. A global mask-guided valid-frame spatio-temporal graph convolution module further enhances robustness to variable-length sequences and differing frame rates, achieving state‑of‑the‑art performance on the E‑Gait dataset and strong generalization across sequence lengths.

By Shirong Lyu, Silu Quan, Yixuan Ding, Chengpeng Wang
arXiv Computer Vision
Sep 3

Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding

The paper introduces MiRA, a plug‑in framework that reweights framewise attention in Vision Transformer video models to better capture subtle facial dynamics for expression recognition. MiRA computes frame‑level confidence and intra‑frame concentration from self‑attention maps, redistributing attention toward localized facial cues without adding trainable parameters. Two modes—an exact post‑softmax redistribution and a lightweight flashLite pre‑softmax approximation—are proposed, and experiments on facial expression recognition benchmarks show consistent gains over strong ViT baselines.

By Seongro Yoon, Donghyeon Cho, Jinsun Park, Fran\c{c}ois Br\'emond
arXiv AI
Sep 7

Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion

The paper introduces a multimodal emotion recognition framework that combines audio and visual feature extraction with an attention-based fusion strategy. Audio features include Wav2Vec2 embeddings, MFCCs, and statistical acoustic descriptors, fused via a BiLSTM, while video features are extracted using a ResNet50-BiLSTM architecture. A multi-head attention mechanism fuses these modalities, and experiments on MELD and IEMOCAP show significant accuracy and robustness gains, especially in unbalanced data settings.

By Xu Lin, Ke Wang, Hui Kang, Xinying Wang
arXiv Computer Vision
Aug 26

Three-Stream Temporal-Shift Attention Network Based on Self-Knowledge Distillation for Micro-Expression Recognition

arXiv:2406.17538v4 Announce Type: replace Abstract: Micro-expressions are subtle facial movements that occur spontaneously when people try to conceal real emotions. Micro-expression recognition is cr...

By Guanghao Zhu, Lin Liu, Yuhao Hu, Haixin Sun, Fang Liu, Xiaohui Du, Ruqian Hao, Juanxiu Liu, Yong Liu, Jing Zhang