arXiv Computer Vision By Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, Anand Bhattad

Principia: Relational Physics Tests for Video Models

Read the original on arXiv Computer Vision →

Principia is a new benchmark that tests video models on Newtonian physics by evaluating relational consistency between paired objects, independent of calibration. It covers eight phenomena—gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass‑spring oscillation—across translational, rotational, collisional, and oscillatory dynamics using real‑world scenes. The benchmark introduces a calibration‑independent consistency score and shows that current state‑of‑the‑art video generators perform poorly on it, with the best vision‑language model achieving only 67% accuracy on detecting physics violations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
1d ago

SnapPhysics: A Physics-Aware Scene Graph from a Single View for Interactive Mixed Reality Scenes

SnapPhysics is a training‑free framework that reconstructs 3D objects and estimates their physical properties—mass, friction, and center of gravity—from a single image. It combines instance‑level 3D reconstruction with a physics‑aware scene graph to provide geometric grounding and inter‑object relationships for vision‑language model reasoning. Experiments on 3D‑FRONT and real captured scenes show significant improvements over existing methods, enabling physically interactive mixed reality experiences without manual tuning.

By Suji Kang, Seok-Young Kim, Young Bin Kim, Taewook Ha, Dieter Schmalstieg, Shohei Mori, Woontack Woo
arXiv AI
Sep 10

CALIPER: Clean Scenes Cannot Rank Physical Inference in Pretrained Visual Representations

CALIPER is a new benchmark that tests whether pretrained visual encoders can infer physical properties such as mass and friction from images. The test involves striking an object twice at known speeds, showing a third strike only up to contact, and asking a linear readout on frozen features to predict how far the object slides. Results show that in clean, fixed‑camera scenes all representations perform similarly, but when camera, lighting, and clutter are varied, only encoders that truly infer physics—like V‑JEPA 2—maintain performance, while random or raw pixel representations fail.

By Aman Mehta, Riya Baviskar
arXiv Computer Vision
Sep 10

MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs

arXiv:2609.09528v1 Announce Type: new Abstract: Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: h...

By Dhairya Bhatia, Bishoy Galoaa, Oliver Fritsche, Shahid Kamal, Muhammad Obaidullah Abdul Salam, Umer Saleem, Om Rastogi, Frania Felix Chettiar, Nesli Erdogmus, Sarah Ostadabbas