arXiv Computer Vision

See the Change, Keep the Flow: Unsupervised Action Segmentation via Spectral-Temporal Representation Learning

Hugging Face Trending Papers
Jun 8

Vision-Language Guided Hyperspectral Object Tracking via Semantics Fusion and Contextual Template Updating

Hyperspectral object tracking (HOT) leverages the rich spectral information provided by hyperspectral videos (HSVs), offering substantial potential for object tracking. However, efficiently extracting and exploiting spectral information from redundant spectral bands remains a fundamental challenge, which severely limits model generalization and tracking performance.

arXiv Computer Vision
Sep 22

Action-Slot: Structured Action-Centric Representation Learning for Multi-Agent Atomic Activity Understanding

The paper introduces Action‑Slot, a structured action‑centric representation learning framework for multi‑agent atomic activity understanding. It reformulates slot attention into activity‑aligned slots, parallel spatio‑temporal updates, and background regularization to disentangle concurrent, asynchronous activities directly from raw video. Additionally, an attention‑difference pseudo‑mask method enables weakly supervised localization, and a new synthetic dataset, TACO, provides balanced atomic activity coverage with pixel‑level annotations.

By Yu-Ho Chang, Chi-Hsi Kung, Yi-Hsuan Tsai, Yi-Ting Chen
Hugging Face Trending Papers
Jul 13

Temporal Feature Distillation for Label-Efficient Precise Event Spotting in Sports Videos

Precise Event Spotting (PES) requires distinguishing visually similar yet semantically distinct adjacent frames, making it fundamentally different from image classification and coarse action recognition. Although self-distillation methods such as DINO have shown strong representation learning ability in images, we find that directly applying them to PES is ineffective: without supervised guidance, subtle but crucial motion cues are often suppressed as noise, leading to representations that are insensitive to precise event boundaries.

arXiv Computer Vision
Aug 27

Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining

The paper introduces Skeleton-Language feature Pooling Switching, a weakly‑supervised vision‑language pretraining strategy for skeleton‑based zero‑shot spatio‑temporal action localization. It replaces video‑level pooling with instance‑level feature computation during inference, enabling the model to estimate unseen actions without costly annotations. Additionally, Scene‑Mixed Discriminative Contrastive Learning is proposed to separate actions at the instance level within mixed scenes using a MIL framework, and experiments on four public datasets confirm the method’s effectiveness.

By Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma, Kazuki Tsutsukawa, Taiki Sekii