Knowing Isn't Always Saying: When Do Spatial Encodings Reach Answers in Vision-Language Models?
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2608.23074v1 Announce Type: new Abstract: Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understoo...
arXiv:2606. 23763v1 Announce Type: cross Abstract: Recent work typically assesses vision--language consistency using attention distributions of answer-side tokens.
Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whether spatial reasoning...
arXiv:2607. 03358v1 Announce Type: cross Abstract: We study how visual information is routed in vision-language models (VLMs).
Vision-language models (VLMs) are increasingly evaluated on complex image and video understanding tasks, yet conventional metrics primarily assess final-answer quality and reveal little about how diff...
arXiv:2609.02000v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly evaluated on complex image and video understanding tasks, yet conventional metrics primarily assess fina...