arXiv AI

Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models

The paper examines how Vision Language Models (VLMs) can automatically extract structured procedural knowledge from industrial troubleshooting guides, which are typically flowchart-like diagrams combining spatial layout and technical language. It evaluates two VLMs using two prompting strategies—standard instruction-guided and an augmented approach that highlights layout patterns—and finds that each model shows different trade-offs between sensitivity to layout and robustness to semantic content. These insights help determine which VLM and prompting method is most suitable for integrating such guides into operator support systems.

arXiv AI
Aug 25

FlowExtract: Procedural Knowledge Extraction from Maintenance Flowcharts

FlowExtract is a pipeline designed to convert ISO 5807-standardized maintenance flowcharts into directed graphs. It separates node detection—using YOLOv8 and EasyOCR—from connectivity reconstruction, employing a novel edge detection method that traces arrowheads back to source nodes. Evaluations on industrial troubleshooting guides show high node detection accuracy and significant improvement over vision‑language model baselines for edge extraction.

By Guillermo Gil de Avalle, Laura Maruster, Eric Sloot, Christos Emmanouilidis
arXiv AI
Aug 25

AI-based worker guidance in assembly and disassembly operations using multimodal ego/exo-centric data capture and structured task knowledge

This paper introduces a data‑centric method that extracts structured task knowledge from expert demonstrations in assembly and disassembly operations. By jointly encoding temporal and multimodal data from egocentric and exocentric video recordings and narration, the approach produces task representations that support procedural documentation and context‑aware worker guidance. Evaluation on a real‑world disassembly case study shows that video‑based representations capture procedural structure and execution context more effectively than static image‑based methods, underscoring the value of egocentric video understanding for repair, training, and circular manufacturing.

By Vivek Chavan, J\"org Kr\"uger
arXiv AI
Sep 21

DiaVLo: Diagnosing Behaviours of Vision-Language Models

DiaVLo is a diagnostic framework for vision‑language models (VLMs) that uses human curation and VLM generation to create specifications of desired and observed behaviours, revealing potential misalignments. It also offers causal estimates to pinpoint the most influential concepts driving VLM behaviour. Experiments on several open‑source VLMs under classification and generation tasks show that DiaVLo’s behaviour labels correlate with model performance and illuminate how VLMs perceive, organise, and prioritise concepts.

By Lorenzo Corti, Jie Yang
Hugging Face Trending Papers
Jul 23

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. KI-VQA involves multiple sub-problems -referring expression understanding, visual grounding, object recognition, knowledge retrieval, and reasoning-yet existing benchmarks typically report only end-task accuracy, obscuring where failures arise.

arXiv AI
Jul 24

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

arXiv:2607. 21155v1 Announce Type: cross Abstract: Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions.

By Hanseok Oh, Parishad BehnamGhader, Benno Krojer, Hyunji Lee, Paul Liang, Siva Reddy, Verna Dankers
arXiv Computer Vision
Aug 25

Investigating Relational Reasoning in VLMs

arXiv:2608.23518v1 Announce Type: new Abstract: Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or...

By Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan, Xander Yap
arXiv AI
Jun 9

Unification of Closed-Open Industrial Detection Scenarios: New Large-Scale Benchmarks,Challenges and Baselines

arXiv:2606. 07953v1 Announce Type: new Abstract: Large-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detection remains challenging due to two fundamental limitations: (i) the scarcity of large-scale industrial datasets that cover diverse defect categories across multiple domains, and (ii) the reliance on manual prompts (points, boxes, masks) that introduce subjective noise and lack text-visual interaction for fine-grained understanding.

By Zekai Zhang, Jinglin Zhang, Qinghui Chen, Gang Li, Da Chen, Shuainan Jing, He Wang, Dagang Li, Cong Liu, Cong Bai, Shengyong Chen