arXiv AI By Guillermo Gil de Avalle, Laura Maruster, Eric Sloot, Christos Emmanouilidis

FlowExtract: Procedural Knowledge Extraction from Maintenance Flowcharts

Read the original on arXiv AI →

FlowExtract is a pipeline designed to convert ISO 5807-standardized maintenance flowcharts into directed graphs. It separates node detection—using YOLOv8 and EasyOCR—from connectivity reconstruction, employing a novel edge detection method that traces arrowheads back to source nodes. Evaluations on industrial troubleshooting guides show high node detection accuracy and significant improvement over vision‑language model baselines for edge extraction.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 25

Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models

The paper examines how Vision Language Models (VLMs) can automatically extract structured procedural knowledge from industrial troubleshooting guides, which are typically flowchart-like diagrams combining spatial layout and technical language. It evaluates two VLMs using two prompting strategies—standard instruction-guided and an augmented approach that highlights layout patterns—and finds that each model shows different trade-offs between sensitivity to layout and robustness to semantic content. These insights help determine which VLM and prompting method is most suitable for integrating such guides into operator support systems.

By Guillermo Gil de Avalle, Laura Maruster, Christos Emmanouilidis
arXiv Computer Vision
Sep 1

TopoAgent: A Structure-Aware Perception-to-Reasoning Framework for Diagram-to-Graph Topology Extraction with Large Vision-Language Models

The paper introduces TopoBench-180, a human‑verified benchmark of 180 structural diagrams with canonical graph annotations, and TopoAgent, a perception‑to‑reasoning framework that extracts graph topology from diagrams using large vision‑language models. TopoAgent combines grounded perception, global structural priors, node inventory construction, local‑to‑global relation reasoning, and consistency enforcement to progressively build the target graph. Experiments demonstrate that TopoAgent surpasses strong baselines, particularly in edge extraction, thereby advancing multimodal structured understanding for diagram‑to‑graph tasks.

By Bangwei Guo, Xujiang Zhao, Yanchi Liu, Wei Cheng, Shengyu Chen, Dongyue Li, Masaharu Morimoto, Takayuki Kuroda, Dimitris Metaxas, Haifeng Chen