SceneTeract: Probing and Improving Agent-Aware Activity Reasoning in 3D Indoor Scenes
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.
arXiv:2603. 09731v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint.
arXiv:2606. 00095v1 Announce Type: cross Abstract: Vision-Language Navigation (VLN) enables embodied agents to reach target locations in unseen environments by following language instructions.
arXiv:2608.30821v1 Announce Type: cross Abstract: Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and e...
arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.
arXiv:2609.19554v1 Announce Type: cross Abstract: Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence,...