ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 08024v1 Announce Type: cross Abstract: Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility.
The paper investigates how visual presentation affects vision‑language models (VLMs) on the SPaRC spatial planning benchmark. By adding lightweight input‑side scaffolds that keep the visual modality but make spatial structure clearer, the authors achieve up to a 34.0‑percentage‑point accuracy boost across multiple VLMs, and an additional 4.6 points when combined with GRPO training. Analyses reveal that these improvements stem mainly from reduced grounding errors, while rule‑based reasoning remains difficult, highlighting visual presentation as a key determinant of whether VLM benchmarks test grounded perception, downstream reasoning, or both.
arXiv:2606. 12550v1 Announce Type: cross Abstract: Open-world mapless navigation from sparse language instructions requires resolving underspecified goals and inferring which environmental cues are relevant for reaching the goal.
arXiv:2608. 20237v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial planning under explicit or previously unseen rule constraints remains underexplored.
arXiv:2606. 04046v1 Announce Type: cross Abstract: In embodied vision-language decision making tasks such as robotic manipulation and navigation, Vision-Language and Vision-Language-Action Models (VLMs & VLAs) are powerful tools with different benefits: VLMs are better at long-term planning, while VLAs are better at reactive control.
arXiv:2608. 16794v1 Announce Type: cross Abstract: Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on incorrectly grounded entities.