arXiv AI By Pedro Orvalho, Guillem Aleny\`a, Felip Many\`a

MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku

Read the original on arXiv AI →

arXiv:2607. 12711v1 Announce Type: new Abstract: Vision--Language Models (VLMs) have recently demonstrated promising performance on structured visual reasoning tasks, including grid-based puzzles.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 7

Think-Verify-Revise: Neuro-Symbolic Visual Reasoning with Vision-Language Models and Dynamic Logic Tensor Networks

The paper introduces a Neuro‑Symbolic framework that integrates a Vision‑Language Model (VLM) for automatic induction of First‑Order Logic (FOL) rules with a Dynamic Logic Tensor Network (D‑LTN) for differentiable rule verification. In a closed iterative loop, the VLM proposes candidate rules (Think), the D‑LTN verifies them against visual embeddings (Verify), and failures guide the VLM to refine its hypotheses (Revise). Evaluated on the ViSudo‑PC benchmark across four visual domains, the system successfully induces Sudoku constraint rules from only three training examples and achieves AUC scores that match or surpass prior methods such as NeuPSL and LTN.

By Homayoun Afshari, Pietro Basci, Alessandro Russo, Lia Morra
arXiv AI
Jun 10

V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions

arXiv:2512. 11995v2 Announce Type: replace-cross Abstract: While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in most benchmarks, they often struggle in practice with complex open-ended tasks, which usually require multiple rounds of exploration and reasoning in the visual space.

By Chenrui Fan, Yijun Liang, Shweta Bhardwaj, Kwesi Cobbina, Ming Li, Tianyi Zhou