D-CLOT: Double Closed Loop Optimal Transport for Unsupervised Action Segmentation
arXiv:2608. 05877v1 Announce Type: cross Abstract: Optimal transport (OT) has emerged as an effective framework for unsupervised action segmentation.
arXiv:2608. 05877v1 Announce Type: cross Abstract: Optimal transport (OT) has emerged as an effective framework for unsupervised action segmentation.
arXiv:2608.29980v1 Announce Type: new Abstract: Unsupervised action segmentation is a challenging task. It involves finding action categories and boundaries in videos without labels. Existing Optimal...
arXiv:2602.05718v2 Announce Type: replace Abstract: Point-supervised Temporal Action Localization (PTAL) adopts a lightly frame-annotated paradigm (\textit{i.e.}, labeling only a single frame per act...
arXiv:2505. 20894v2 Announce Type: replace Abstract: Despite recognized limitations in modeling long-range temporal dependencies, Human Activity Recognition (HAR) has traditionally relied on a sliding window approach to segment labeled datasets.
arXiv:2606. 19408v1 Announce Type: new Abstract: Latent actions provide a compact interface between action-free video and downstream decision-making, yet existing Latent Action Models (LAMs) force every transition through a fixed-capacity bottleneck.
Hyperspectral object tracking (HOT) leverages the rich spectral information provided by hyperspectral videos (HSVs), offering substantial potential for object tracking. However, efficiently extracting and exploiting spectral information from redundant spectral bands remains a fundamental challenge, which severely limits model generalization and tracking performance.
arXiv:2608.27562v1 Announce Type: new Abstract: Translating continuous, noisy egocentric video streams into discrete, temporally ordered action steps is fraught with visual challenges. Heavy ego-moti...
The paper introduces Action‑Slot, a structured action‑centric representation learning framework for multi‑agent atomic activity understanding. It reformulates slot attention into activity‑aligned slots, parallel spatio‑temporal updates, and background regularization to disentangle concurrent, asynchronous activities directly from raw video. Additionally, an attention‑difference pseudo‑mask method enables weakly supervised localization, and a new synthetic dataset, TACO, provides balanced atomic activity coverage with pixel‑level annotations.
Precise Event Spotting (PES) requires distinguishing visually similar yet semantically distinct adjacent frames, making it fundamentally different from image classification and coarse action recognition. Although self-distillation methods such as DINO have shown strong representation learning ability in images, we find that directly applying them to PES is ineffective: without supervised guidance, subtle but crucial motion cues are often suppressed as noise, leading to representations that are insensitive to precise event boundaries.
arXiv:2607. 04750v1 Announce Type: new Abstract: We present FM-ChangeNet, a pathwise-supervised framework for change detection that reformulates bi-temporal reasoning as continuous transport in feature space rather than static endpoint comparison.
The paper introduces Skeleton-Language feature Pooling Switching, a weakly‑supervised vision‑language pretraining strategy for skeleton‑based zero‑shot spatio‑temporal action localization. It replaces video‑level pooling with instance‑level feature computation during inference, enabling the model to estimate unseen actions without costly annotations. Additionally, Scene‑Mixed Discriminative Contrastive Learning is proposed to separate actions at the instance level within mixed scenes using a MIL framework, and experiments on four public datasets confirm the method’s effectiveness.
Fine-grained understanding of operating room (OR) activity could enable workflow-aware assistance, yet remains difficult due to clutter, occlusions, and limited sensing. The prevailing approach to model this environment is scene graphs as an interpretable representation of OR interactions.