FlowExtract is a pipeline designed to convert ISO 5807-standardized maintenance flowcharts into directed graphs. It separates node detection—using YOLOv8 and EasyOCR—from connectivity reconstruction, employing a novel edge detection method that traces arrowheads back to source nodes. Evaluations on industrial troubleshooting guides show high node detection accuracy and significant improvement over vision‑language model baselines for edge extraction.
By Guillermo Gil de Avalle, Laura Maruster, Eric Sloot, Christos Emmanouilidis
arXiv:2608.22174v1 Announce Type: new
Abstract: Unified multimodal models (UMMs) can perform both understanding and generation, raising a central question: can visual generation improve understanding...
By Yubo Zhu, Zhehan Kan, Jingyi Yang, Miaolin Chen, Jinbo Xing, Kai Zhu, Zijian Wang, Sheng Zhong, Wei Tong
arXiv:2607. 13035v1 Announce Type: cross Abstract: Cloud services experience frequent incidents that require rapid diagnosis and resolution.
By Srihari Unnikrishnan, Jaskaran Singh Walia, Drishti Goel, Supriyo Ghosh
This paper introduces a data‑centric method that extracts structured task knowledge from expert demonstrations in assembly and disassembly operations. By jointly encoding temporal and multimodal data from egocentric and exocentric video recordings and narration, the approach produces task representations that support procedural documentation and context‑aware worker guidance. Evaluation on a real‑world disassembly case study shows that video‑based representations capture procedural structure and execution context more effectively than static image‑based methods, underscoring the value of egocentric video understanding for repair, training, and circular manufacturing.
By Vivek Chavan, J\"org Kr\"uger
arXiv:2606. 12969v1 Announce Type: new Abstract: The power distribution network is critical to reliable electricity delivery, yet traditional inspection methods face limitations in semantic understanding, generalization, and closed-loop automation.
By Quan Quan
arXiv:2606. 17904v1 Announce Type: new Abstract: Language models increasingly serve as advisory systems in maintenance operations.
By Guillermo Gil de Avalle, Laura Maruster, Shaina Raza, Christos Emmanouilidis
DiaVLo is a diagnostic framework for vision‑language models (VLMs) that uses human curation and VLM generation to create specifications of desired and observed behaviours, revealing potential misalignments. It also offers causal estimates to pinpoint the most influential concepts driving VLM behaviour. Experiments on several open‑source VLMs under classification and generation tasks show that DiaVLo’s behaviour labels correlate with model performance and illuminate how VLMs perceive, organise, and prioritise concepts.
By Lorenzo Corti, Jie Yang
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. KI-VQA involves multiple sub-problems -referring expression understanding, visual grounding, object recognition, knowledge retrieval, and reasoning-yet existing benchmarks typically report only end-task accuracy, obscuring where failures arise.
arXiv:2607. 15418v1 Announce Type: new Abstract: We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings -- a core media in architecture, civil, and many other engineering practices.
By Yoonhwa Jung, Junryu Fu, Mani Golparvar-Fard
arXiv:2607. 21155v1 Announce Type: cross Abstract: Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions.
By Hanseok Oh, Parishad BehnamGhader, Benno Krojer, Hyunji Lee, Paul Liang, Siva Reddy, Verna Dankers
arXiv:2608.23518v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or...
By Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan, Xander Yap
arXiv:2606. 07953v1 Announce Type: new Abstract: Large-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detection remains challenging due to two fundamental limitations: (i) the scarcity of large-scale industrial datasets that cover diverse defect categories across multiple domains, and (ii) the reliance on manual prompts (points, boxes, masks) that introduce subjective noise and lack text-visual interaction for fine-grained understanding.
By Zekai Zhang, Jinglin Zhang, Qinghui Chen, Gang Li, Da Chen, Shuainan Jing, He Wang, Dagang Li, Cong Liu, Cong Bai, Shengyong Chen