Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2508. 10956v3 Announce Type: replace-cross Abstract: Inspired by human categorization, visual reasoning about object properties, such as physical attributes and functions, involves identifying and recognizing low-level details and higher-level abstractions.
arXiv:2608. 01664v1 Announce Type: cross Abstract: We present our ImageCLEF 2026 Multimodal Reasoning system for the Visual Multiple Choice Question Answering (Visual MCQ) and Visual Open Question Answering (Visual OpenQA) subtasks.
arXiv:2512. 11995v2 Announce Type: replace-cross Abstract: While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in most benchmarks, they often struggle in practice with complex open-ended tasks, which usually require multiple rounds of exploration and reasoning in the visual space.
The paper introduces Illusion-Reasoning, a benchmark that uses real-world visual illusion images paired with annotated question-answer pairs to evaluate Large Vision Language Models (LVLMs). It argues that current evaluations focus too narrowly on perception or specific domains, and that a joint assessment of perception and reasoning in open-world settings is lacking. Using this benchmark, the authors demonstrate that many LVLMs’ reasoning abilities are less advanced than previously claimed, offering new insights and directions for optimization.
arXiv:2607. 00491v1 Announce Type: cross Abstract: Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input.
arXiv:2608. 13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).