arXiv Computer Vision

PSEE: Progressive Sensor Event Expansion for Point-Supervised Temporal Action Localization

The paper introduces PSEE, a method for progressive sensor event expansion that generates pseudo action segments from point-supervised labels in wearable sensor streams. By leveraging semantic activations, sensor-specific transition evidence, and adaptive temporal ownership, PSEE produces high‑quality pseudo boundaries that can train standard temporal action localization detectors without altering their inference. Experiments on four inertial‑sensing benchmarks show that PSEE outperforms adapted point‑supervised baselines, works with various detectors, and remains robust to different point sampling strategies.

arXiv Computer Vision
Aug 27

Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining

The paper introduces Skeleton-Language feature Pooling Switching, a weakly‑supervised vision‑language pretraining strategy for skeleton‑based zero‑shot spatio‑temporal action localization. It replaces video‑level pooling with instance‑level feature computation during inference, enabling the model to estimate unseen actions without costly annotations. Additionally, Scene‑Mixed Discriminative Contrastive Learning is proposed to separate actions at the instance level within mixed scenes using a MIL framework, and experiments on four public datasets confirm the method’s effectiveness.

By Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma, Kazuki Tsutsukawa, Taiki Sekii
arXiv AI
Sep 18

PointEvent: Rethinking Event-based Tiny Object Detection via Serialized Motion Evidence Accumulation

PointEvent introduces serialized motion evidence accumulation for event-based tiny object detection, treating motion continuity as an ordered evidence propagation process. The method organizes event streams into locality‑preserving spatiotemporal paths and chronology‑preserving temporal paths, alternating serialized scans across complementary orders to consolidate fragmented motion evidence. A lightweight event‑wise state‑space framework with a high‑resolution event branch and compact context modulation achieves state‑of‑the‑art performance with the fewest parameters and fastest inference among compared methods.

By Zongze Wu, Baofeng Jia, Weiqi Yan, Jingyuan Zhang, Yu Zang, Xiaoyu Chen, Jing Han
arXiv Computer Vision
Sep 7

FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models

The paper introduces FailureSpot, a label‑efficient method for detecting failures at the timestamp level in vision‑language‑action (VLA) policies. It first generates weak supervision from unlabeled VLA action chunks by identifying abnormal patterns, then employs active learning to annotate only the most uncertain trajectories. Experiments on multiple VLA policies demonstrate improved performance for both timestamp‑level and trajectory‑level failure detection.

By Jie Ma, Zongxi Liu, Yi Zhu
arXiv Computer Vision
Aug 26

ConsensusTAS: Self-Supervised Temporal Action Segmentation for Long-Horizon Construction Videos

ConsensusTAS is a self‑supervised, label‑free method for temporal action segmentation in long construction videos. It segments continuous video streams into distinct activity phases by leveraging internal consensus among candidate segmentations, and it has been evaluated on three public datasets, outperforming state‑of‑the‑art methods. In real‑world construction footage, the model successfully identified fine‑grained actions within bricklaying, and it can run on a CPU, making it suitable for mobile robotic platforms.

By Xiaoshan Zhou, Yafei Sun
Hugging Face Trending Papers
Jul 13

Temporal Feature Distillation for Label-Efficient Precise Event Spotting in Sports Videos

Precise Event Spotting (PES) requires distinguishing visually similar yet semantically distinct adjacent frames, making it fundamentally different from image classification and coarse action recognition. Although self-distillation methods such as DINO have shown strong representation learning ability in images, we find that directly applying them to PES is ineffective: without supervised guidance, subtle but crucial motion cues are often suppressed as noise, leading to representations that are insensitive to precise event boundaries.

arXiv Machine Learning
Jul 14

Towards Real-World Wearable Motion Reconstruction

arXiv:2607. 09780v1 Announce Type: cross Abstract: The modern-day surge in popularity of wearable devices poses a fundamentally unique motion capture problem: reconstructing full-body movement from any set of sensing hardware worn at a given moment.

By Andrea Boscolo Camiletto, Rishabh Dabral, Eduardo Alvarado, Thabo Beeler, Marc Habermann, Christian Theobalt