The paper examines how Vision Language Models (VLMs) can automatically extract structured procedural knowledge from industrial troubleshooting guides, which are typically flowchart-like diagrams combining spatial layout and technical language. It evaluates two VLMs using two prompting strategies—standard instruction-guided and an augmented approach that highlights layout patterns—and finds that each model shows different trade-offs between sensitivity to layout and robustness to semantic content. These insights help determine which VLM and prompting method is most suitable for integrating such guides into operator support systems.
By Guillermo Gil de Avalle, Laura Maruster, Christos Emmanouilidis
arXiv:2602.11678v2 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual understanding, yet they suffer from a critical limitation: struct...
By Chengwei Ma, Zhen Tian, Zhou Zhou, Zhixian Xu, Xiaowei Zhu, Xia Hua, Si Shi, F. Richard Yu
The paper introduces TopoBench-180, a human‑verified benchmark of 180 structural diagrams with canonical graph annotations, and TopoAgent, a perception‑to‑reasoning framework that extracts graph topology from diagrams using large vision‑language models. TopoAgent combines grounded perception, global structural priors, node inventory construction, local‑to‑global relation reasoning, and consistency enforcement to progressively build the target graph. Experiments demonstrate that TopoAgent surpasses strong baselines, particularly in edge extraction, thereby advancing multimodal structured understanding for diagram‑to‑graph tasks.
By Bangwei Guo, Xujiang Zhao, Yanchi Liu, Wei Cheng, Shengyu Chen, Dongyue Li, Masaharu Morimoto, Takayuki Kuroda, Dimitris Metaxas, Haifeng Chen
arXiv:2602.13880v2 Announce Type: replace
Abstract: Graph property detection aims to determine whether a graph exhibits certain structural properties, such as being Hamiltonian. Recently, learning-ba...
By Jiahao Xie, Guangmo Tong
arXiv:2607. 05841v1 Announce Type: cross Abstract: Structured representation can characterize semantic objects and relationships in images.
By Zhiguang Zhou, Fengling Zheng, Miaoxin Hu, Lina You, Jin Wen, Huan Liu, Wei Zhang, Dekun Qian, Yuhua Liu, Wei Chen, Yigang Wang, Yong Wang
arXiv:2608.21825v1 Announce Type: new
Abstract: Learning adjacency matrices from node-link images is a fundamental problem for recovering structured graph information from visual observations. Existi...
By Jiahao Xie, Guangmo Tong