arXiv AI By Pulkit Madan, Sanjay Haresh, Reza Ebrahimi, Sunny Panchal, Apratim Bhattacharyya, Roland Memisevic

On Locality and Length Generalization in Visual Reasoning

Read the original on arXiv AI →

arXiv:2607. 09061v1 Announce Type: cross Abstract: A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single global computation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

A Progressive Training Strategy for Embodied Vision-Language Models to Mitigate Spatio-Temporal Hallucinations

The paper introduces a progressive training strategy for embodied vision‑language models aimed at reducing spatio‑temporal hallucinations. It first creates a Chain‑of‑Thought dataset that breaks complex reasoning into detailed spatiotemporal steps, then uses supervised pre‑training on this dataset followed by fine‑tuning with weakly‑labeled data. Experiments show the method improves backbone accuracy and narrows the forward‑backward performance gap from over 70% to 6.53%, indicating stronger dynamic reasoning and fewer temporal biases.

By Xiaoda Yang, Shuai Yang, Can Wang, Jingyang Xue, Menglan Tang, Checheng Yu, Xunzhe Zhou, Sashuai Zhou, Tao Jin, Lixin Yang, Xiangyu Yue, Zhou Zhao
arXiv Computer Vision
Sep 1

Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting

The paper introduces structured video prompting, a training‑free inference‑time technique that augments input videos with lightweight spatial and temporal structure to provide explicit anchors for evidence organization. By applying this method to two video benchmarks and two open video‑language models, the authors demonstrate performance improvements across several tasks, with gains varying by model and task. The study suggests that failures in video‑language models stem not only from reasoning capacity but also from how video evidence is presented during inference.

By Sadegh Mohammadian