SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
Read the original on arXiv Computation and Language →SpaceCast-Bench is a new benchmark that evaluates predictive spatial reasoning in vision‑language models, moving beyond simple spatial perception to tasks that require constructing scenes, anticipating interventions, and reasoning about unseen outcomes. It contains 3,862 questions from 182 real‑world scenes across 16 task types and three difficulty levels—static perception, local prediction, and global prediction—testing scene understanding, spatial state updating, and relational inference. Evaluation of 21 models shows a large performance gap, with the best model achieving only 58.0% versus 87.2% human accuracy, and fine‑tuning on the benchmark’s programmatically generated data can substantially improve model performance.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.