arXiv AI By Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He

VISTA: A Visual Harness for Reasoning in an Interactive World

Read the original on arXiv AI →

The paper introduces VISTA, a visual harness that equips a general-purpose multimodal model with long‑horizon vision and a lossless visual memory. VISTA enables the model to directly perceive and actively retrieve past observations, allowing it to reorganize visual input during reasoning. On the ARC‑AGI‑3 benchmark, VISTA boosts Claude Opus 5.0’s Relative Human Action Efficiency from 40.68 to a perfect 100.00, completing all 25 public games with 57.4% fewer actions than first‑time human participants, and it also outperforms baselines on three additional visual game and puzzle benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 14

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

arXiv:2607. 11436v1 Announce Type: new Abstract: Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening evidence-grounded reasoning.

By Wencheng Ye, Yi Bin, Yujuan Ding, Hongye Fang, Zheng Wang, Xing Xu, Jingkuan Song, Yun Zhang, Sirui Da, Heng Tao Shen
arXiv AI
Jun 16

Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning

arXiv:2606. 15231v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated impressive capabilities in many visual tasks, but they often struggle with factual grounding when confronted with complex, open-world scenarios.

By Zhengbo Zhang, Changtao Miao, Jinbo Su, Zhaowen Zhou, Chunxia Zhang, Xukai Wang, Ruiqi Liu, Kaiyuan Zheng, Jiansheng Cai, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan