arXiv Computer Vision
The paper introduces HML-FSAR, a hierarchical metric learning framework for few-shot action recognition. It incorporates a spatial‑enhanced module, temporal MHA, heterogeneous alignment, spatial‑temporal fusion, and dictionary learning to build a comprehensive feature pipeline. Progressive constraints—center, alignment, contrastive, dictionary, and prototype metrics—are applied from frame‑level representations to final prototypes, improving feature compactness, alignment, discriminability, and robustness.
arXiv:2504.10079v5 Announce Type: replace
Abstract: Few-shot action recognition (FSAR) aims to recognize novel action categories with few exemplars. Existing methods typically learn frame-level repre...
By Hongyu Qu, Ling Xing, Jiachao Zhang, Rui Yan, Yazhou Yao, Xiangbo Shu
arXiv:2508. 12745v2 Announce Type: replace-cross Abstract: Image set classification (ISC), which can be viewed as a task of comparing similarities between sets consisting of unordered heterogeneous images with variable quantities and qualities, has attracted growing research attention in recent years.
By Xizhan Gao, Wei Hu
The paper introduces Skeleton-Language feature Pooling Switching, a weakly‑supervised vision‑language pretraining strategy for skeleton‑based zero‑shot spatio‑temporal action localization. It replaces video‑level pooling with instance‑level feature computation during inference, enabling the model to estimate unseen actions without costly annotations. Additionally, Scene‑Mixed Discriminative Contrastive Learning is proposed to separate actions at the instance level within mixed scenes using a MIL framework, and experiments on four public datasets confirm the method’s effectiveness.
By Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma, Kazuki Tsutsukawa, Taiki Sekii
arXiv:2609.28086v1 Announce Type: cross
Abstract: We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings...
By Sandra Arcos-Holzinger, Debashish Chakraborty, Rohita Mocharla, Will Walden, Andrew Yates, Reno Kriz, Sarah M. Erfani, James Bailey, Vishal M. Patel, Sanjeev Khudanpur
arXiv:2609.21522v1 Announce Type: new
Abstract: Recent pre-trained foundation models provide rich multi-modal priors for downstream 3D vision tasks. However, the effectiveness of these representation...
By Hang Cheng, Yan Chen, Mingyu Fan, Long Zeng
Whole slide image (WSI) classification is an evidence-driven task, where diagnostic cues are often sparse, spatially organized, and class-dependent. Existing MIL and vision-language methods aggregate a large pool of patch features into a single global slide representation.
Action Quality Assessment (AQA) aims to objectively evaluate performance quality from action videos. Most existing methods follow a ``one-by-one'' paradigm, training a separate model for each action type.
arXiv:2602.05718v2 Announce Type: replace
Abstract: Point-supervised Temporal Action Localization (PTAL) adopts a lightly frame-annotated paradigm (\textit{i.e.}, labeling only a single frame per act...
By Yunchuan Ma, Laiyun Qing, Guorong Li, Yuqing Liu, Yuankai Qi, Qingming Huang
arXiv:2604. 03919v2 Announce Type: replace-cross Abstract: We present the first systematic study of Sparse Autoencoders (SAEs) on video representations.
By Atahan Dokme, Sriram Vishwanath
Visible-infrared person re-identification (VI-ReID) suffers from cross-modal discrepancies and limited discriminative capabilities, leading to suboptimal recognition performance. Current approaches ex...
arXiv:2607. 03131v1 Announce Type: cross Abstract: Modern video surveillance systems generate far more video streams than human operators can effectively monitor, making automated analysis essential for timely detection of security events.
By Estera Dumitru, Stelian Sp\^inu
arXiv:2608.29186v1 Announce Type: new
Abstract: Federated Video Domain Adaptation (FVDA) enables collaborative learning across distributed and non-IID video datasets while preserving privacy, but is...
By Lee En-Yi Hannah, Haozhi Cao, Yuecong Xu