arXiv Computer Vision

Few-Shot Video Recognition via Hierarchical Metric Learning

The paper introduces HML-FSAR, a hierarchical metric learning framework for few-shot action recognition. It incorporates a spatial‑enhanced module, temporal MHA, heterogeneous alignment, spatial‑temporal fusion, and dictionary learning to build a comprehensive feature pipeline. Progressive constraints—center, alignment, contrastive, dictionary, and prototype metrics—are applied from frame‑level representations to final prototypes, improving feature compactness, alignment, discriminability, and robustness.

arXiv Computer Vision
Aug 27

Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining

The paper introduces Skeleton-Language feature Pooling Switching, a weakly‑supervised vision‑language pretraining strategy for skeleton‑based zero‑shot spatio‑temporal action localization. It replaces video‑level pooling with instance‑level feature computation during inference, enabling the model to estimate unseen actions without costly annotations. Additionally, Scene‑Mixed Discriminative Contrastive Learning is proposed to separate actions at the instance level within mixed scenes using a MIL framework, and experiments on four public datasets confirm the method’s effectiveness.

By Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma, Kazuki Tsutsukawa, Taiki Sekii
arXiv AI
Sep 24

LAYERSCOPE: A Layerwise Characterization of Video and Multimodal Learned Representations

arXiv:2609.28086v1 Announce Type: cross Abstract: We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings...

By Sandra Arcos-Holzinger, Debashish Chakraborty, Rohita Mocharla, Will Walden, Andrew Yates, Reno Kriz, Sarah M. Erfani, James Bailey, Vishal M. Patel, Sanjeev Khudanpur