arXiv AI By Renjie Gu, Kaichen Zhou, Yan Luo, Mengyu Wang

GeoWorld-VLM: Geometry from World Models for Vision-Language Models

Read the original on arXiv AI →

arXiv:2605. 16713v2 Announce Type: replace-cross Abstract: Modern Vision-Language Models (VLMs) achieve strong semantic recognition, yet remain brittle on elementary spatial relations such as left of, on, behind, and between.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos

PhysMLLMs introduces physics-inspired spatial continuity priors into video multimodal large language models to address spatio‑temporal inconsistencies such as jitter, drift, and identity switches. The method, called Global Representation Prior Alignment (REPA‑Global), distills global visual representations from a frozen DINOv2 teacher during training, aligning student representations without affecting inference speed. Experiments on multiple video benchmarks show improved segmentation mask quality and cross‑frame consistency, especially for challenging scenarios involving small targets, fast motion, occlusion, and distractors, while maintaining comparable performance on single‑frame image segmentation and general VLM tasks.

By Siyao Yan, Bo Han, Jisheng Dang, Bimei Wang, Shude Wang, Hong Peng, Yulan Guo, Jianhuang Lai, Bin Hu, Tat-SengChua
arXiv Computation and Language
4d ago

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

arXiv:2609.38177v1 Announce Type: cross Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...

By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong