arXiv AI

STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets

STATERA is a method that adapts a pretrained video backbone with mostly frozen weights and a lightweight temporal tubelet mixer to estimate the center-of-mass (CoM) of opaque, asymmetric rigid bodies from short monocular videos. It introduces the HiddenMass Benchmark, consisting of 50K simulated MuJoCo trajectories and a 63-sequence real-world test set with calibrated CoM ground truth. In simulation, STATERA reduces normalized CoM error from 41.7% to 25.2%, and in zero-shot sim-to-real transfer, its phase‑aware variant consistently predicts movement toward the true hidden offset, improving physics capture from 2.6% to 41.0%.

arXiv Computer Vision
Aug 31

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

The paper introduces Parallel Tube Decoding (PTD), a generative approach for spatio‑temporal video grounding that splits the task into a temporal block and simultaneous time‑conditioned spatial blocks, eliminating token‑level and trajectory‑level dependencies. PTD uses Decoupled Block Attention to allow parallel spatial generation while maintaining shared video‑query context, and incorporates localization‑aware policy optimization for temporal boundaries and spatial geometry. Experiments on VidSTG show PTD cuts tube completion latency by 79× and boosts spatial decoding throughput by 92× compared to autoregressive decoding, while improving grounding accuracy and performing well on related tasks such as temporal grounding, VideoQA, and referring video object tracking.

By Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang, Fahad Shahbaz Khan, Salman Khan
arXiv AI
Jun 30

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

arXiv:2606. 29531v1 Announce Type: cross Abstract: We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs.

By Weisong Liu, Haochen Wang, Kuan Gao, Yuhao Wang, Yikang Zhou, Zhongwei Ren, Jacky Mai, Anna Wang, Yanwei Li, Jason Li, Zhaoxiang Zhang
arXiv AI
Sep 15

LPA-CWM: A Learned Physical Adjudicator for Motion Reasoning with Counterfactual World Models

LPA-CWM introduces a Learned Physical Adjudicator (LPA) to improve counterfactual world models (CWM) for motion reasoning by learning to weight candidate responses based on visual context and response structure. The 3.0M‑parameter LPA is trained on dense MOVi‑F trajectories while keeping the CWM predictor and intervention generator frozen. A new Completeness‑aware Motion Correspondence (CMC) protocol evaluates localization, trajectory completeness, visibility, and continuity, and LPA‑CWM achieves significant gains on DAVIS and Kinetics subsets.

By Kunwei Wu, Xiang Liu, Guocai Yao, Junming Chen, Zhikang Chen, Min Zhang, Pengwei Wang, Sen Cui
arXiv Computer Vision
6d ago

Reliability-Regulated Trajectory Optimization for Progressive COLMAP-Free 3D Gaussian Splatting

The paper introduces a reliability-regulated trajectory optimization framework for progressive COLMAP‑free 3D Gaussian Splatting (3DGS). It uses a self‑supervised bidirectional cycle‑consistency mechanism to control camera trajectory estimation through forward motion propagation and retrospective trajectory correction, thereby reducing error compounding without external priors. Experiments on Tanks and Temples and CO3D‑V2 demonstrate improved camera trajectory accuracy and novel‑view rendering quality compared to existing unposed baselines.

By Zijian Wu, Jinliang Wang, Zidian Lin, Ying Song, Ziqian Lu, Hanjie Ma, Zhen Ye, Mingfeng Jiang
Hugging Face Trending Papers
Jul 21

Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.