arXiv AI By Patrick Rim, Tom Long, Ekta Prashnani, Ruth Rosenholtz, Ben Boudaoud, Peter Xenopoulos, Alex Wong, Joohwan Kim, Jae-Hyun Jung

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests

Read the original on arXiv AI →

arXiv:2607. 22864v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.