arXiv AI

Visual Graph Scaffolds for Structural Reasoning in Large Language Models

arXiv:2606. 02673v1 Announce Type: new Abstract: Graphs have been used to enhance large language models (LLMs) for structured reasoning, mostly as external knowledge sources are provided to models at test time.

arXiv Machine Learning
Sep 21

VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering

VISPATH is a visual‑intent‑guided path reasoning framework designed for multimodal knowledge graph question answering (MM‑KGQA). It first identifies a reliable starting entity by fusing multimodal grounding with graph‑structural cues, then iteratively discovers and refines reasoning paths using hop‑specific multimodal intent and a reasoning‑chain pruning step. The framework is evaluated on the newly introduced VISPATH‑Bench, which tests two‑to‑four‑hop reasoning, and demonstrates consistent improvements over strong baselines, even surpassing GPT‑5.4 when using GPT‑4o as the backbone.

By Jinke Wu, Zhengpin Li, Mengzhe Jia, Yang Li, Wentao Zhang
arXiv AI
Aug 5

CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning

arXiv:2608. 02833v1 Announce Type: cross Abstract: Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains.

By Xuehang Guo, Pingyue Zhang, Ruiyi Zhang, Zhenhailong Wang, Hanrui Lyu, Heng Ji, Tong Sun, Qingyun Wang, Manling Li
arXiv AI
Aug 7

ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion

arXiv:2608. 05833v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images.

By Jiafan Li, Mengxue Yang, Jiaqi Zhu, Liang Chang, Ying Li, Hongan Wang
arXiv Machine Learning
Sep 4

When Vision Meets Graphs: A Survey on Graph Reasoning and Learning

The survey titled "When Vision Meets Graphs: A Survey on Graph Reasoning and Learning" reviews how visual depictions of graphs can be used as inputs for graph reasoning and learning. It highlights that while Graph Neural Networks dominate graph machine learning, most pipelines ignore the visual form of graphs, despite scientists routinely interpreting graphs visually. The paper organizes existing work into three threads—vision for graph reasoning, vision for graph learning, and scientific graphs—aiming to clarify current capabilities and chart a path toward foundation models that perceive and reason about graphs like scientists do.

By Xinjian Zhao, Wei Pang, Zhixuan Yu, Xiangru Jian, Xiaozhuang Song, Yaoyao Xu, Zhongkai Xue, Dingshuo Chen, Shu Wu, Philip Torr, Tianshu Yu
arXiv Machine Learning
Sep 18

VisKG-LM: Compiling Knowledge Graphs into Visual Memory for Multiple-Choice Question Answering

VisKG‑LM proposes compiling retrieved knowledge graph subgraphs into static visual memories rather than re‑encoding them during each inference step. The method serializes each subgraph as Relation‑Labeled Paths, renders them as images that preserve the graph’s branching structure, and caches these images for reuse. At inference, a language model processes the question and candidate text first, then consults the cached visual memory only at its final layer, yielding improved performance on CommonsenseQA, OpenBookQA, and MedQA‑USMLE compared to both text‑only baselines and a large vision‑language model.

By Yixin Peng, Er Jin, Shiwei Luo, Diego Collarana, Stefan Decker
arXiv AI
Jun 10

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering

arXiv:2510. 04514v3 Announce Type: replace Abstract: Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts-those requiring precise visual interpretation rather than relying on textual shortcuts.

By Rachneet Kaur, Nishan Srishankar, Zhen Zeng, Sumitra Ganesh, Manuela Veloso