In this paper, we present the solution developed by our team, XInsight Lab, which achieved first place in Track 3 of the 4th EI-MIGA-IJCAI Challenge with a test accuracy of 0. 76923.
arXiv:2607. 20820v1 Announce Type: new Abstract: Body-based emotion recognition is important for real-time affective systems, but graph-based skeleton models can be computationally expensive.
By Christian Arzate Cruz, Stefanos Gkikas, Houshyar Asadi
The paper introduces GaitMoE, an action‑detection based mixture‑of‑experts framework for occluded gait recognition, leveraging temporal and action experts to infer missing body parts from adjacent frames and gait cycles. It also presents a new Occluded Gait database (OccGait) with diverse occlusion scenarios and annotations, and demonstrates superior performance on OccGait, OccCASIA‑B, Gait3D, and GREW datasets.
By Panjian Huang, Yunjie Peng, Saihui Hou, Chunshui Cao, Xu Liu, Zhiqiang He, Yongzhen Huang
The paper introduces a multimodal emotion recognition framework that combines audio and visual feature extraction with an attention-based fusion strategy. Audio features include Wav2Vec2 embeddings, MFCCs, and statistical acoustic descriptors, fused via a BiLSTM, while video features are extracted using a ResNet50-BiLSTM architecture. A multi-head attention mechanism fuses these modalities, and experiments on MELD and IEMOCAP show significant accuracy and robustness gains, especially in unbalanced data settings.
By Xu Lin, Ke Wang, Hui Kang, Xinying Wang
arXiv:2609.08038v2 Announce Type: cross
Abstract: Smart healthcare monitoring systems require precise action recognition to ensure well-being and timely intervention in critical situations such as fa...
By Diwas Lamsal, Pramod Wickramatilake, Jednipat Moonrinta, Mongkol Ekpanyapong, Matthew N. Dailey
The paper introduces Skeleton-Language feature Pooling Switching, a weakly‑supervised vision‑language pretraining strategy for skeleton‑based zero‑shot spatio‑temporal action localization. It replaces video‑level pooling with instance‑level feature computation during inference, enabling the model to estimate unseen actions without costly annotations. Additionally, Scene‑Mixed Discriminative Contrastive Learning is proposed to separate actions at the instance level within mixed scenes using a MIL framework, and experiments on four public datasets confirm the method’s effectiveness.
By Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma, Kazuki Tsutsukawa, Taiki Sekii
arXiv:2606. 28377v1 Announce Type: cross Abstract: HAR using Inertial Measurement Unit (IMU) sensors is vital for healthcare monitoring and rehabilitation.
By Saeid Arabzadeh, Farshad Almasganj, Mohammad Mahdi Ahmadi
MTF‑Net is a Multi‑Modal Temporal Feature Fusion Network that jointly models kinematic, appearance, and contextual cues for pedestrian intention prediction. It fuses four modalities—bounding‑box dynamics, human pose keypoints, local context, and scene‑level semantics—within a recurrent framework enhanced by gated linear units (GLUs) and an attention‑guided fusion head. Evaluations on the PIE and JAAD benchmarks show that MTF‑Net outperforms recent transformer‑ and graph‑based models, achieving up to 0.95 AUC on PIE and 0.94 AUC on JAAD while maintaining real‑time performance.
By Md Mahfuzur Rahman, Pengzhan Zhou, A. F. M. Abdun Noor, Md Imam Ahasan, Md Mustafizur Rahman, Fang Qu
Fine-grained understanding of operating room (OR) activity could enable workflow-aware assistance, yet remains difficult due to clutter, occlusions, and limited sensing. The prevailing approach to model this environment is scene graphs as an interpretable representation of OR interactions.
arXiv:2607. 10984v1 Announce Type: cross Abstract: Existing Stochastic 3D Human Motion Prediction models are fundamentally constrained by hard-coding the skeleton kinematics, severely limiting generalization, preventing cross-dataset training, and requiring complex data retargeting.
By Cecilia Curreli, Florian Hofherr, Dominik Muhle, Abhishek Saroha, Riccardo Marin, Daniel Cremers
arXiv:2607. 11839v1 Announce Type: cross Abstract: This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthcare-oriented training environments.
By Divya Mereddy, Jeevan Beedareddy
arXiv:2607. 00716v1 Announce Type: cross Abstract: Skeleton-based action recognition has achieved remarkable success by exploiting joint coordinates and their topological connections, yet prevailing methods overwhelmingly assume complete and clean skeleton inputs.
By Yingjie Dai, Tianyang Xu, Yanglin Deng, Xiao-Jun Wu, Josef Kittler