arXiv AI By Guillermo Gil de Avalle, Laura Maruster, Christos Emmanouilidis

Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models

Read the original on arXiv AI →

The paper examines how Vision Language Models (VLMs) can automatically extract structured procedural knowledge from industrial troubleshooting guides, which are typically flowchart-like diagrams combining spatial layout and technical language. It evaluates two VLMs using two prompting strategies—standard instruction-guided and an augmented approach that highlights layout patterns—and finds that each model shows different trade-offs between sensitivity to layout and robustness to semantic content. These insights help determine which VLM and prompting method is most suitable for integrating such guides into operator support systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 25

FlowExtract: Procedural Knowledge Extraction from Maintenance Flowcharts

FlowExtract is a pipeline designed to convert ISO 5807-standardized maintenance flowcharts into directed graphs. It separates node detection—using YOLOv8 and EasyOCR—from connectivity reconstruction, employing a novel edge detection method that traces arrowheads back to source nodes. Evaluations on industrial troubleshooting guides show high node detection accuracy and significant improvement over vision‑language model baselines for edge extraction.

By Guillermo Gil de Avalle, Laura Maruster, Eric Sloot, Christos Emmanouilidis
arXiv AI
Aug 25

AI-based worker guidance in assembly and disassembly operations using multimodal ego/exo-centric data capture and structured task knowledge

This paper introduces a data‑centric method that extracts structured task knowledge from expert demonstrations in assembly and disassembly operations. By jointly encoding temporal and multimodal data from egocentric and exocentric video recordings and narration, the approach produces task representations that support procedural documentation and context‑aware worker guidance. Evaluation on a real‑world disassembly case study shows that video‑based representations capture procedural structure and execution context more effectively than static image‑based methods, underscoring the value of egocentric video understanding for repair, training, and circular manufacturing.

By Vivek Chavan, J\"org Kr\"uger