arXiv AI By Quan Quan

Multi-Modal Agents for Power Distribution Defect Detection: An Evaluation of Foundation Models

Read the original on arXiv AI →

arXiv:2606. 12969v1 Announce Type: new Abstract: The power distribution network is critical to reliable electricity delivery, yet traditional inspection methods face limitations in semantic understanding, generalization, and closed-loop automation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 9

Unification of Closed-Open Industrial Detection Scenarios: New Large-Scale Benchmarks,Challenges and Baselines

arXiv:2606. 07953v1 Announce Type: new Abstract: Large-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detection remains challenging due to two fundamental limitations: (i) the scarcity of large-scale industrial datasets that cover diverse defect categories across multiple domains, and (ii) the reliance on manual prompts (points, boxes, masks) that introduce subjective noise and lack text-visual interaction for fine-grained understanding.

By Zekai Zhang, Jinglin Zhang, Qinghui Chen, Gang Li, Da Chen, Shuainan Jing, He Wang, Dagang Li, Cong Liu, Cong Bai, Shengyong Chen
arXiv AI
Sep 1

AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance

arXiv:2506.03828v4 Announce Type: replace Abstract: AI for Industrial Asset Lifecycle Management aims to automate complex operational workflows, such as condition monitoring and maintenance schedulin...

By Dhaval Patel, Shuxin Lin, James Rayfield, Nianjun Zhou, Chathurangi Shyalika, Suryanarayana R Yarrabothula, Roman Vaculin, Natalia Martinez, Fearghal O'donncha, Jayant Kalagnanam
arXiv AI
Aug 25

Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models

The paper examines how Vision Language Models (VLMs) can automatically extract structured procedural knowledge from industrial troubleshooting guides, which are typically flowchart-like diagrams combining spatial layout and technical language. It evaluates two VLMs using two prompting strategies—standard instruction-guided and an augmented approach that highlights layout patterns—and finds that each model shows different trade-offs between sensitivity to layout and robustness to semantic content. These insights help determine which VLM and prompting method is most suitable for integrating such guides into operator support systems.

By Guillermo Gil de Avalle, Laura Maruster, Christos Emmanouilidis