arXiv AI

Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning

arXiv:2607. 23077v1 Announce Type: new Abstract: Modeling multi-entity temporal data requires capturing dependencies across entities, time, and their interactions.

arXiv AI
Jun 29

EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning

arXiv:2603. 09731v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint.

By Chengjun Yu, Xuhan Zhu, Chaoqun Du, Pengfei Yu, Wei Zhai, Yang Cao, Zheng-Jun Zha