arXiv Machine Learning

The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show

arXiv:2606. 05328v1 Announce Type: cross Abstract: Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simulators.

arXiv AI
Jun 9

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis

arXiv:2606. 09646v1 Announce Type: cross Abstract: We study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how this information varies across model families, layers, and probe types.

By Samuele Punzo, Niccol\`o Caselli, Ippokratis Pantelidis, Francesco Massafra, Salvatore Lo Sardo, Mohammadreza Salehi