Motion-Saliency Complementary Masked Modeling for Point Cloud Video Understanding
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608.24093v1 Announce Type: cross Abstract: Self-supervised representation learning for 4D point cloud videos is challenging because annotations are costly and reconstruction-based pretraining...
Video object segmentation (VOS) is a fundamental task in video understanding, requiring accurate delineation and consistent tracking of objects across frames. While supervised methods achieve strong performance, they rely on densely annotated datasets that are costly to obtain and have limited domain coverage.
arXiv:2609.40347v1 Announce Type: new Abstract: We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of r...
arXiv:2606. 06853v1 Announce Type: cross Abstract: The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding.
arXiv:2608.30450v1 Announce Type: new Abstract: Video virtual try-on aims to transfer a target garment onto a moving person across video frames. Current methods rely on human parsing masks or pose ke...
arXiv:2608.30388v1 Announce Type: cross Abstract: Cross-view video representation learning aims to capture viewpoint-invariant action semantics despite substantial appearance changes across egocentri...