arXiv AI By Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat

On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes

Read the original on arXiv AI →

arXiv:2410. 19553v2 Announce Type: replace-cross Abstract: This paper explores the impact of occlusions in video action detection.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
3d ago

FOMO: Forget the Concept, Don't Miss Out on the Scene in Selective Video Unlearning

FOMO is a training‑based selective video unlearning method that prioritizes preserving the original scene while removing targeted concepts. It localizes concept‑related representations for modification and employs a preservation mechanism that maintains non‑target scene information without auxiliary data. The approach extends to motion unlearning, enabling removal of concepts defined by temporal behavior, and achieves a strong balance between concept removal and scene preservation.

By {\L}ukasz Rudnik, Agnieszka Polowczyk, Alicja Polowczyk, Przemys{\l}aw Spurek
arXiv Computer Vision
Sep 17

Occluded Gait Recognition with Mixture of Experts: An Action Detection Perspective

The paper introduces GaitMoE, an action‑detection based mixture‑of‑experts framework for occluded gait recognition, leveraging temporal and action experts to infer missing body parts from adjacent frames and gait cycles. It also presents a new Occluded Gait database (OccGait) with diverse occlusion scenarios and annotations, and demonstrates superior performance on OccGait, OccCASIA‑B, Gait3D, and GREW datasets.

By Panjian Huang, Yunjie Peng, Saihui Hou, Chunshui Cao, Xu Liu, Zhiqiang He, Yongzhen Huang
arXiv AI
Aug 28

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

LeVJEPA is a video encoder that eliminates the need for architectural asymmetries, exponential-moving-average target encoders, stop-gradients, and capacity-limited predictors used in prior self‑supervised methods. It trains a single encoder with an invariance loss over global and local views, regularized by SIGReg to prevent collapse, and achieves strong performance with far less pretraining compute. The approach also allows block‑causal attention, making temporal ordering a property of the encoder itself, and matches or surpasses state‑of‑the‑art baselines on both appearance‑centric and motion‑centric benchmarks.

By Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner