arXiv AI

STAG: Spatio-temporal Evolving Structural Representation of Action Units for Micro-expression Recognition

arXiv:2606. 28083v1 Announce Type: cross Abstract: Micro-expression recognition is challenging due to subtle and short-lived facial muscle movements.

arXiv Computer Vision
Aug 26

Three-Stream Temporal-Shift Attention Network Based on Self-Knowledge Distillation for Micro-Expression Recognition

arXiv:2406.17538v4 Announce Type: replace Abstract: Micro-expressions are subtle facial movements that occur spontaneously when people try to conceal real emotions. Micro-expression recognition is cr...

By Guanghao Zhu, Lin Liu, Yuhao Hu, Haixin Sun, Fang Liu, Xiaohui Du, Ruqian Hao, Juanxiu Liu, Yong Liu, Jing Zhang
arXiv AI
Aug 13

HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation

arXiv:2608. 12187v1 Announce Type: cross Abstract: Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spatial and temporal reasoning as separate stages, which may weaken unified spatial-temporal interdependencies inherent in human motion and compress frame-level structural information before temporal modelling.

By Ruochen Li, Shuang Chen, Wenke E, Farshad Arvin, Amir Atapour-Abarghouei
arXiv Computer Vision
Sep 3

Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding

The paper introduces MiRA, a plug‑in framework that reweights framewise attention in Vision Transformer video models to better capture subtle facial dynamics for expression recognition. MiRA computes frame‑level confidence and intra‑frame concentration from self‑attention maps, redistributing attention toward localized facial cues without adding trainable parameters. Two modes—an exact post‑softmax redistribution and a lightweight flashLite pre‑softmax approximation—are proposed, and experiments on facial expression recognition benchmarks show consistent gains over strong ViT baselines.

By Seongro Yoon, Donghyeon Cho, Jinsun Park, Fran\c{c}ois Br\'emond
arXiv Computer Vision
Sep 11

Single-Stream Multi-Feature Fusion with Temporal Robustness for Gait Emotion Recognition

The paper introduces SV-GCN, a single-stream multi-feature fusion framework for 3D skeleton-based gait emotion recognition that incorporates temporal invariance. It uses intra-frame relative motion features to remove frame-rate sensitivity and embeds heterogeneous cues at shallow layers for early fusion, avoiding multi-stream complexity. A global mask-guided valid-frame spatio-temporal graph convolution module further enhances robustness to variable-length sequences and differing frame rates, achieving state‑of‑the‑art performance on the E‑Gait dataset and strong generalization across sequence lengths.

By Shirong Lyu, Silu Quan, Yixuan Ding, Chengpeng Wang
arXiv AI
Jul 8

HST-HGN: Heterogeneous Spatial-Temporal Hypergraph Networks with Bidirectional State Space Models for Global Fatigue Assessment

arXiv:2604. 08435v2 Announce Type: replace-cross Abstract: It remains challenging to assess driver fatigue from untrimmed videos under constrained computational budgets, due to the difficulty of modeling long-range temporal dependencies in subtle facial expressions.

By Changdao Chen, Qinqiuhong Ye, Hao Chen, Jinyu Wang