arXiv Machine Learning By Luca M. Schulze Buschoff, Konstantinos Voudouris, Can Demircan, Eric Schulz

Can Vision Language Models Learn Intuitive Physics from Interaction?

Read the original on arXiv Machine Learning →

arXiv:2602. 06033v2 Announce Type: replace Abstract: Pre-trained vision language models do not have good intuitions about the physical world.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 14

PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning

arXiv:2607. 10190v1 Announce Type: cross Abstract: Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and fundamental physical principles is essential.

By Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu