arXiv:2602.05718v2 Announce Type: replace
Abstract: Point-supervised Temporal Action Localization (PTAL) adopts a lightly frame-annotated paradigm (\textit{i.e.}, labeling only a single frame per act...
By Yunchuan Ma, Laiyun Qing, Guorong Li, Yuqing Liu, Yuankai Qi, Qingming Huang
The paper introduces Skeleton-Language feature Pooling Switching, a weakly‑supervised vision‑language pretraining strategy for skeleton‑based zero‑shot spatio‑temporal action localization. It replaces video‑level pooling with instance‑level feature computation during inference, enabling the model to estimate unseen actions without costly annotations. Additionally, Scene‑Mixed Discriminative Contrastive Learning is proposed to separate actions at the instance level within mixed scenes using a MIL framework, and experiments on four public datasets confirm the method’s effectiveness.
By Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma, Kazuki Tsutsukawa, Taiki Sekii
arXiv:2606. 00260v1 Announce Type: cross Abstract: Human Activity Recognition (HAR) from ambient sensors enables smart-home applications such as health monitoring and assisted living.
By Zishuai Liu, Ruili Fang, Jin Lu, Fei Dou
PointEvent introduces serialized motion evidence accumulation for event-based tiny object detection, treating motion continuity as an ordered evidence propagation process. The method organizes event streams into locality‑preserving spatiotemporal paths and chronology‑preserving temporal paths, alternating serialized scans across complementary orders to consolidate fragmented motion evidence. A lightweight event‑wise state‑space framework with a high‑resolution event branch and compact context modulation achieves state‑of‑the‑art performance with the fewest parameters and fastest inference among compared methods.
By Zongze Wu, Baofeng Jia, Weiqi Yan, Jingyuan Zhang, Yu Zang, Xiaoyu Chen, Jing Han
arXiv:2608. 15861v1 Announce Type: new Abstract: Fine-grained wrist activity recognition can support applications such as procedural step guidance and context-aware assistance, yet acquiring labeled data for every new task, user, and activity granularity remains a bottleneck.
By Aidan Bradshaw, Riku Arakawa, Xin Liu, Karan Ahuja
arXiv:2505. 20894v2 Announce Type: replace Abstract: Despite recognized limitations in modeling long-range temporal dependencies, Human Activity Recognition (HAR) has traditionally relied on a sliding window approach to segment labeled datasets.
By Marius Bock, Juergen Gall, Michael Moeller, Kristof Van Laerhoven
The paper introduces FailureSpot, a label‑efficient method for detecting failures at the timestamp level in vision‑language‑action (VLA) policies. It first generates weak supervision from unlabeled VLA action chunks by identifying abnormal patterns, then employs active learning to annotate only the most uncertain trajectories. Experiments on multiple VLA policies demonstrate improved performance for both timestamp‑level and trajectory‑level failure detection.
By Jie Ma, Zongxi Liu, Yi Zhu
ConsensusTAS is a self‑supervised, label‑free method for temporal action segmentation in long construction videos. It segments continuous video streams into distinct activity phases by leveraging internal consensus among candidate segmentations, and it has been evaluated on three public datasets, outperforming state‑of‑the‑art methods. In real‑world construction footage, the model successfully identified fine‑grained actions within bricklaying, and it can run on a CPU, making it suitable for mobile robotic platforms.
By Xiaoshan Zhou, Yafei Sun
arXiv:2604. 00767v2 Announce Type: replace Abstract: Wearable human activity recognition (HAR) has made steady progress, yet much of this progress remains grounded in fixed-window, closed-set classification benchmarks.
By Lala Shakti Swarup Ray, Mengxi Liu, Alcina Pinto, Deepika Gurung, Daniel Geissler, Paul Lukowoicz, Bo Zhou
Precise Event Spotting (PES) requires distinguishing visually similar yet semantically distinct adjacent frames, making it fundamentally different from image classification and coarse action recognition. Although self-distillation methods such as DINO have shown strong representation learning ability in images, we find that directly applying them to PES is ineffective: without supervised guidance, subtle but crucial motion cues are often suppressed as noise, leading to representations that are insensitive to precise event boundaries.
arXiv:2607. 09780v1 Announce Type: cross Abstract: The modern-day surge in popularity of wearable devices poses a fundamentally unique motion capture problem: reconstructing full-body movement from any set of sensing hardware worn at a given moment.
By Andrea Boscolo Camiletto, Rishabh Dabral, Eduardo Alvarado, Thabo Beeler, Marc Habermann, Christian Theobalt
arXiv:2604. 18064v2 Announce Type: replace Abstract: Human motion world models should capture motion's intentionality by being executable: adaptable to different actions and capable of assessing motion quality.
By Rimvydas Rubavicius, Manisha Dubey, N. Siddharth, Subramanian Ramamoorthy