arXiv AI

HST-HGN: Heterogeneous Spatial-Temporal Hypergraph Networks with Bidirectional State Space Models for Global Fatigue Assessment

arXiv:2604. 08435v2 Announce Type: replace-cross Abstract: It remains challenging to assess driver fatigue from untrimmed videos under constrained computational budgets, due to the difficulty of modeling long-range temporal dependencies in subtle facial expressions.

arXiv AI
Aug 13

HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation

arXiv:2608. 12187v1 Announce Type: cross Abstract: Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spatial and temporal reasoning as separate stages, which may weaken unified spatial-temporal interdependencies inherent in human motion and compress frame-level structural information before temporal modelling.

By Ruochen Li, Shuang Chen, Wenke E, Farshad Arvin, Amir Atapour-Abarghouei
Hugging Face Trending Papers
Jul 1

Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach

Skeleton-based action recognition has achieved remarkable success by exploiting joint coordinates and their topological connections, yet prevailing methods overwhelmingly assume complete and clean skeleton inputs. In real-world deployments, such as egocentric vision, crowded surveillance, wearable devices, or edge robotics, limited field-of-view (FoV) frequently causes substantial joint visibility dropout, leading to severe performance degradation that existing models are largely unprepared to handle.

arXiv Computer Vision
Aug 27

Bidirectional Temporal Dynamics Modeling for EEG-based Driving Fatigue Recognition

The paper introduces DeltaGateNet, a framework that models bidirectional temporal dynamics in EEG signals for driving fatigue recognition. It uses a Bidirectional Delta module to separate positive and negative temporal differences, and a Gated Temporal Convolution module to capture long‑term dependencies while preserving channel specificity. Experiments on SEED‑VIG and SADT datasets show that DeltaGateNet outperforms existing methods, achieving high intra‑subject and inter‑subject accuracies across balanced and unbalanced data.

By Yip Tin Po, Jianming Wang, Yutao Miao, Jiayan Zhang, Yunxu Zhao, Xiaomin Ouyang, Zhihong Li, Nevin L. Zhang
arXiv AI
Sep 25

UNWIND: Any-Length Facial Video for Stress Detection without Temporal Windowing

UNWIND is a facial‑video framework that detects stress by treating an entire recording as a single input, avoiding the need for temporal windowing or segmentation. It folds the video’s temporal dimension into the channel dimension of a 2‑D spatial representation and processes it with an asymmetric‑attention architecture. Experiments on a 58‑subject stress dataset show that using all 3,600 frames (stride τ = 1) yields a 69.73 % accuracy, comparable to the best 70.02 % accuracy at τ = 15, while computational cost varies from 12.48 to 348.78 GFLOPs.

By Stefanos Gkikas, Christian Arzate Cruz, Eric Nichols, Giorgos Giannakakis, Randy Gomez
arXiv Computer Vision
Sep 16

Hyper-RED: Scalable Event Pre-training via Semantic Hypergraph Distillation

Hyper-RED introduces a scalable image-to-event pretraining framework that transfers high‑order semantic structures via hypergraphs, avoiding rigid pixel‑wise alignment. By constructing image, event, and cross‑modal hypergraphs and applying a hypergraph relational distillation loss, the method preserves local relational consistency and event‑specific characteristics while inheriting image‑derived semantic organization. Experiments across five event datasets show consistent scaling from ViT‑S to ViT‑L and state‑of‑the‑art performance.

By Meisen Wang, Zhiqiang Tian, Wei Bao, Chengjie Wang, Shaoyi Du, Siqi Li
arXiv AI
Aug 12

FUSE: Frame-Unified Stress Estimation from Facial Video

arXiv:2608. 10442v1 Announce Type: cross Abstract: Automatic stress detection from facial video offers a practical path to non-intrusive affect monitoring, yet existing video-based approaches commonly decompose full recordings into short temporal windows before classification.

By Stefanos Gkikas, Thomas Kassiotis, Yang Guo, Guangliang Li, Giorgos Giannakakis
arXiv Computer Vision
Aug 26

Stack Transformer Based Spatial-Temporal Attention Model for Dynamic Sign Language and Fingerspelling Recognition

The paper introduces the Sequential Spatio-Temporal Attention Network (SSTAN), a Transformer-based architecture that replaces traditional Graph Convolutional Networks for sign language recognition. SSTAN uses a hierarchical, stacked design with Spatial Multi-Head Attention to model joint relationships within frames and Temporal Multi-Head Attention to capture long-range dependencies across frames, eliminating the need for predefined skeletal graphs. Experiments on large-scale datasets (WLASL, JSL, KSL) show that SSTAN, trained from scratch, achieves state‑of‑the‑art performance in fingerspelling categories and outperforms other skeleton‑only methods on WLASL, highlighting its data efficiency and ability to learn complex spatio‑temporal patterns.

By Koki Hirooka, Abu Saleh Musa Miah, Tatsuya Murakami, Md. Al Mehedi Hasan, Yong Seok Hwang, Jungpil Shin