Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition
arXiv:2608. 05115v1 Announce Type: cross Abstract: Can computer vision help make classrooms safer?
Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom incident recognition from CCTV-style observations.
arXiv:2608. 05115v1 Announce Type: cross Abstract: Can computer vision help make classrooms safer?
arXiv:2604. 03401v4 Announce Type: replace-cross Abstract: Understanding student engagement usually requires time-consuming manual observation or invasive recording that raises privacy concerns.
The paper introduces CamVLM, a framework that equips large vision‑language models with the ability to actively control camera viewpoints for improved surveillance video understanding. It presents two new datasets: CCTV‑Anomaly, a large‑scale surveillance video collection with detailed captions and event annotations, and CamTrack‑53K, an object‑centric viewpoint trajectory dataset for learning camera actions. Using reinforcement learning, CamVLM learns long‑horizon observation strategies, achieving state‑of‑the‑art performance in both passive and dynamic viewpoint settings.
MINT is a foundation model that directly predicts world-space two-hand trajectories from egocentric RGB video, jointly estimating camera motion, hand states, and hand presence in a single spatiotemporal representation. It uses an open-source labeling pipeline, EGOPIPELINE, to generate large-scale pseudo-labels for pretraining, followed by fine-tuning on a small set of high-quality joint annotations. The model outperforms existing multi-stage approaches in accuracy and speed, and generalizes zero‑shot to unseen egocentric datasets.
arXiv:2605.21957v2 Announce Type: replace Abstract: Video anomaly detection is critical for public safety and security, yet remains highly challenging despite extensive research due to large variatio...
Track2Art is a motion‑centric framework that recovers articulated object models from RGB‑D interaction videos by lifting 2D point tracks into 3D trajectories. It groups these trajectories into rigid‑part hypotheses and uses learned‑analytic reasoning to infer directed kinematic relations, joint types, and joint geometry. On the PartNet‑Mobility benchmark, it achieves 0.695 Point IoU and 0.410 end‑to‑end J@20 without requiring ground‑truth part counts or test‑time optimization.
arXiv:2603. 09731v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint.
The paper presents a generative framework that estimates category-level 6D pose and 3D size of objects from a single RGB image, using score-based diffusion models to produce a multi-hypothesis pose distribution. It replaces costly likelihood pruning with a Mean Shift approach to isolate the mode as the final pose estimate, achieving state-of-the-art results on the REAL275 benchmark. The method also decouples detection from pose estimation, enabling robust zero-shot generalisation on the Wild6D dataset and extending naturally to video sequences by propagating the pose distribution over time.
arXiv:2605.17610v2 Announce Type: replace-cross Abstract: The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-wo...
arXiv:2609.36937v1 Announce Type: cross Abstract: Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation,...
arXiv:2608. 10932v1 Announce Type: cross Abstract: Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation.
The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.