Hierarchical Relation-augmented Representation Generalization for Few-shot Action Recognition
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper introduces HML-FSAR, a hierarchical metric learning framework for few-shot action recognition. It incorporates a spatial‑enhanced module, temporal MHA, heterogeneous alignment, spatial‑temporal fusion, and dictionary learning to build a comprehensive feature pipeline. Progressive constraints—center, alignment, contrastive, dictionary, and prototype metrics—are applied from frame‑level representations to final prototypes, improving feature compactness, alignment, discriminability, and robustness.
arXiv:2401. 10805v4 Announce Type: replace-cross Abstract: We introduce the novel concept of visually Connecting Actions and Their Effects (CATE) in video understanding.
Action Quality Assessment (AQA) aims to objectively evaluate performance quality from action videos. Most existing methods follow a ``one-by-one'' paradigm, training a separate model for each action type.
arXiv:2602.05718v2 Announce Type: replace Abstract: Point-supervised Temporal Action Localization (PTAL) adopts a lightly frame-annotated paradigm (\textit{i.e.}, labeling only a single frame per act...
arXiv:2609.09736v1 Announce Type: new Abstract: Video Temporal Grounding (VTG) localizes the video segment that matches a natural-language query. Many queries describe an action performed by a partic...
Fine-grained understanding of operating room (OR) activity could enable workflow-aware assistance, yet remains difficult due to clutter, occlusions, and limited sensing. The prevailing approach to model this environment is scene graphs as an interpretable representation of OR interactions.