arXiv Machine Learning

LaGSplat: Inferring Physics-Governed Interactive Simulation from Monocular Video Using Latent Lagrangian Gaussian Splatting

arXiv:2608. 16324v1 Announce Type: cross Abstract: We present LaGSplat (Latent Lagrangian Gaussian Splatting), a framework that infers interactive, physics-governed dynamics from one or a few monocular videos.

arXiv AI
2d ago

STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets

STATERA is a method that adapts a pretrained video backbone with mostly frozen weights and a lightweight temporal tubelet mixer to estimate the center-of-mass (CoM) of opaque, asymmetric rigid bodies from short monocular videos. It introduces the HiddenMass Benchmark, consisting of 50K simulated MuJoCo trajectories and a 63-sequence real-world test set with calibrated CoM ground truth. In simulation, STATERA reduces normalized CoM error from 41.7% to 25.2%, and in zero-shot sim-to-real transfer, its phase‑aware variant consistently predicts movement toward the true hidden offset, improving physics capture from 2.6% to 41.0%.

By Animesh Varma
arXiv Computer Vision
Aug 27

4DStreamCtrl: Interactive Video Generation with Online 4D Control

The paper introduces 4DStreamCtrl, a system that unifies camera motion, object trajectories, and depth into a single 3D point‑track representation, enabling joint control, depth editing, and motion transfer in a single forward pass. By mining in‑the‑wild video for 3D motion supervision and encoding it with a lightweight Geometric Motion Head, the authors train a causal streaming student that can generate arbitrarily long videos in just four denoising steps, achieving 20 FPS on a single high‑end GPU for 480p video. This approach outperforms prior camera‑only, 2D, and offline‑3D methods in motion‑control precision while maintaining temporal coherence over hundreds of frames, thereby enabling interactive 4D‑controllable streaming generation for the first time.

By Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu
Hugging Face Trending Papers
Jul 21

Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.

arXiv Machine Learning
Aug 27

JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.

By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
arXiv AI
Aug 28

KnockGS:interaction-Grounded Calibrationof Physical Gaussian Representations

The paper introduces KnockGS, a framework that calibrates the elasticity and density of 3D Gaussian objects by analyzing their dynamic response to known forces. By extracting temporal response features from observed dynamics, KnockGS estimates material scales, freezes them into the simulator, and demonstrates that these estimates remain accurate for unseen interactions. Evaluation shows that KnockGS outperforms alternative methods in parameter recovery and response fidelity across diverse material targets.

By Chenchen Ge, Hanwen Shen, Bowen Jing, Jiyuan Cai, Xiaofeng Wang, Hongsen Lei, Weitao Zhou, Dandan Zhang, Haibao Yu
arXiv Computer Vision
Sep 14

Physics-Aware Video Generation via Agentic Planning and Graph-Guided Optimization

PhysPlan is a training‑free guidance framework that enhances video diffusion models by incorporating physical awareness through agentic physics simulation. It uses a vision‑language model to generate a Chain‑of‑Visual‑Thought representation of kinematic trajectories and 3D depth, which then drives an object‑centric test‑time optimization that isolates kinematic changes and locks the passive environment. The framework also employs Kinetic Intensity Profiling to adapt hyperparameters to varying physical deformations, and demonstrates superior performance on PhyGenBench and Physics‑IQ benchmarks compared to existing VDM baselines.

By Minh-Loi Nguyen, Xuan-Vu Le, Thanh-Toan Do, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le