← Back to all news
arXiv Computer Vision September 1, 2026 By Wei Wang, Yiding Sun, Yuyan Wang, Zhuoyue Zhang, Zhengqiao Li, Dongfu Yin, Chen Li

Motion-Saliency Complementary Masked Modeling for Point Cloud Video Understanding

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • computer-vision

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Aug 26

Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos

arXiv:2608.24093v1 Announce Type: cross Abstract: Self-supervised representation learning for 4D point cloud videos is challenging because annotations are costly and reconstruction-based pretraining...

By Jheng-Ling Lee, Shang-Tse Chen
ragfine-tuningbenchmarks
More like this →
Hugging Face Trending Papers
Jul 8

`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation

Video object segmentation (VOS) is a fundamental task in video understanding, requiring accurate delineation and consistent tracking of objects across frames. While supervised methods achieve strong performance, they rely on densely annotated datasets that are costly to obtain and have limited domain coverage.

llmscomputer-visionsafety
More like this →
arXiv Computer Vision
1d ago

Image Classifiers are Efficient Self-Supervised Video Representation Learners

arXiv:2609.40347v1 Announce Type: new Abstract: We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of r...

By Owais Iqbal, Sudipta Sarkar, Shyam Marjit, Omprakash Chakraborty, Anirban Chakraborty, Abir Das
llmsragcomputer-visionbenchmarks
More like this →
arXiv AI
Jun 8

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models

arXiv:2606. 06853v1 Announce Type: cross Abstract: The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding.

By Yifan Xu, Chao Zhang, Ruifei Ma, Fei Gao, Zhifei Yang, Jiaxing Qi, Zhipeng Chen
llmsagentsdiffusionmultimodalbenchmarkssafety
More like this →
arXiv Computer Vision
Sep 1

FlowVVTON: Flow-Guided Mask-Free Video Virtual Try-On

arXiv:2608.30450v1 Announce Type: new Abstract: Video virtual try-on aims to transfer a target garment onto a moving person across video frames. Current methods rely on human parsing masks or pose ke...

By Shengyao Chen, Xianbing Sun, Liqing Zhang, Jianfu Zhang
computer-visionsafety
More like this →
arXiv AI
Sep 1

PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning

arXiv:2608.30388v1 Announce Type: cross Abstract: Cross-view video representation learning aims to capture viewpoint-invariant action semantics despite substantial appearance changes across egocentri...

By Youngchae Chee, Hosu Lee, Sungjune Park, Junho Kim, Yong Man Ro
ragbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea