arXiv:2508. 12448v2 Announce Type: replace-cross Abstract: In-context learning (ICL) lets large language models (LLMs) solve new tasks from prompts alone, across an ever-widening range of domains, yet the mechanisms underlying this ability remain poorly understood.
By Yeongwoo Song, Jaeyong Bae, Dong-Kyum Kim, Hawoong Jeong
arXiv:2607. 06522v1 Announce Type: new Abstract: Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tasks and environments.
By Han-Jun Ko, Jr-Jen Chen, Haobo Yuan, Hsin-Ying Lee, Tiancheng Shen, Ming-Hsuan Yang, Yu-Chiang Frank Wang
arXiv:2607. 10190v1 Announce Type: cross Abstract: Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and fundamental physical principles is essential.
By Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu
arXiv:2509. 12263v3 Announce Type: replace Abstract: Large multimodal models (LMMs) encode physical laws observed during training, such as momentum conservation, as parametric knowledge.
By Gautam Sreekumar, Vishnu Naresh Boddeti
arXiv:2603. 07109v2 Announce Type: replace Abstract: Understanding physical transformations is fundamental for reasoning in dynamic environments.
By Dezhi Luo, Yijiang Li, Maijunxian Wang, Tianwei Zhao, Bingyang Wang, Siheng Wang, Pinyuan Feng, Pooyan Rahmanzadehgervi, Ziqiao Ma, Hokin Deng
arXiv:2607. 23899v1 Announce Type: cross Abstract: This exploratory study examines whether a large multimodal language model, GPT-5.
By Roberto Spinelli, Thiago C. Martins
Synthesizing realistic Human-Object Interactions (HOI) is critical for creating embodied avatars and functional virtual environments. However, current data-driven approaches primarily rely on motion capture datasets, which are expensive to scale and limited in functional diversity.
Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on general video understanding tasks, current video-language models still struggle to reliably determine whether an observed event conforms to specific physical laws.
arXiv:2608. 02150v2 Announce Type: replace-cross Abstract: Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities.
By Zhongjie Ba, Shengwang Xu, Peng Cheng, Jinyang Zou, Ting Yu, Zhibo Wang, Zhan Qin
arXiv:2608. 09876v1 Announce Type: cross Abstract: Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must strictly conform to real-world execution dynamics.
By Yapeng Liu, Yuanzhao Zhai, Bo Ding, Huaimin Wang, Lin Wang
arXiv:2510. 14828v3 Announce Type: replace Abstract: Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks successfully.
By Jinrui Liu, Bingyan Nie, Boyu Li, Yaran Chen, Yuze Wang, Shunsen He, Haoran Li
arXiv:2605. 28865v2 Announce Type: replace-cross Abstract: What does a world model learn from physical exploration, without any linguistic supervision?
By Jiayi Fang