arXiv:2608. 14569v1 Announce Type: new Abstract: Neural solvers for constraint satisfaction problems have achieved remarkable in-distribution accuracy, yet they suffer from a fundamental limitation persistent constraint violations occur under distribution shifts even when the model reports high confidence.
By Shufeng Kong, Xiaochuan Zhang, Caihua Liu
arXiv:2606. 03269v1 Announce Type: new Abstract: Visual Question Answering (VQA) is the task of answering questions about images, requiring the integration of multimodal input and reasoning.
By Thomas Eiter, Nelson Higuera Ruiz, Johannes Oetsch
The paper introduces a Neuro‑Symbolic framework that integrates a Vision‑Language Model (VLM) for automatic induction of First‑Order Logic (FOL) rules with a Dynamic Logic Tensor Network (D‑LTN) for differentiable rule verification. In a closed iterative loop, the VLM proposes candidate rules (Think), the D‑LTN verifies them against visual embeddings (Verify), and failures guide the VLM to refine its hypotheses (Revise). Evaluated on the ViSudo‑PC benchmark across four visual domains, the system successfully induces Sudoku constraint rules from only three training examples and achieves AUC scores that match or surpass prior methods such as NeuPSL and LTN.
By Homayoun Afshari, Pietro Basci, Alessandro Russo, Lia Morra
arXiv:2603. 23867v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have been applied to a wide range of reasoning tasks, yet it remains unclear whether they can reason robustly under distribution shifts.
By Weixin Chen, Antonio Vergari, Han Zhao
arXiv:2512. 11995v2 Announce Type: replace-cross Abstract: While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in most benchmarks, they often struggle in practice with complex open-ended tasks, which usually require multiple rounds of exploration and reasoning in the visual space.
By Chenrui Fan, Yijun Liang, Shweta Bhardwaj, Kwesi Cobbina, Ming Li, Tianyi Zhou
arXiv:2607. 16727v1 Announce Type: new Abstract: Autoregressive multimodal large language models (MLLMs) suffer from error snowballing: a single incorrect inference early in a chainof-thought (CoT) trace corrupts all downstream reasoning.
By Zehua Cheng, Wei Dai, Jiahao Sun
arXiv:2607. 02959v1 Announce Type: cross Abstract: We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process.
By Harsh Goel, S P Sharan, Sahil Shah, Minkyu Choi, Joungbin An, Kristen Grauman, Sandeep P. Chinchali
arXiv:2606. 27926v1 Announce Type: new Abstract: Geometry Problem Solving have increasingly adopt the neuro-symbolic paradigm, combining neural intuition with symbolic rigor.
By Can Li, Ting Zhang, Junbo Zhao, Hua Huang
arXiv:2602. 14065v2 Announce Type: replace Abstract: Knowledge-intensive Visual Question Answering (KI-VQA) frequently suffers from severe knowledge conflicts caused by the inherent limitations of open-domain retrieval.
By Kai Ye, Xianwei Mao, Sheng Zhou, Zirui Shao, Ye Mo, Liangliang Liu, Haikuan Huang, Bin Li, Jiajun Bu
The paper introduces a multimodal reasoning framework for cross‑domain visual question answering in Printed Circuit Board Assembly (PCBA) inspection, converting diverse data sources into a unified instruction format and generating verified reasoning traces. It proposes Task‑Aware Group Relative Policy Optimization (GRPO) that uses semantic, distance‑aware, and format rewards to improve choice‑based and counting tasks beyond exact‑match supervision. During inference, the system applies semantic consistency correction, self‑consistency voting, and multi‑model arbitration, achieving an overall score of 83.24 on the PCBA Standard‑to‑Real Grand Challenge leaderboard.
By Jia Li, Li Dai, Peng Jia, Zhenzhen Hu, Chee Seng Chan, Bingkun Bao, Richang Hong
arXiv:2608. 20237v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial planning under explicit or previously unseen rule constraints remains underexplored.
By Yu Chen, Ting Lei, Yaoyi Li, Jia Cai, Zhecen Wu, Yang Liu
arXiv:2609.35942v1 Announce Type: new
Abstract: Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs i...
By Ting-Chih Chen, Emile van Krieken, Shujian Yu, Filip Ilievski