Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline
arXiv:2606. 07965v1 Announce Type: new Abstract: Large Visual Language Models (LVLMs) have achieved remarkable success in vision tasks.
arXiv:2606. 07953v1 Announce Type: new Abstract: Large-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detection remains challenging due to two fundamental limitations: (i) the scarcity of large-scale industrial datasets that cover diverse defect categories across multiple domains, and (ii) the reliance on manual prompts (points, boxes, masks) that introduce subjective noise and lack text-visual interaction for fine-grained understanding.
arXiv:2606. 07965v1 Announce Type: new Abstract: Large Visual Language Models (LVLMs) have achieved remarkable success in vision tasks.
arXiv:2608.29783v1 Announce Type: new Abstract: Industrial anomaly detection is a critical component of modern manufacturing. Most traditional unsupervised methods rely on modelling normal feature di...
CF-YOLO introduces a real‑time detection framework for camouflaged micro‑defects on industrial components, combining a Context‑Perception Aggregation Module (CPAM) that fuses large‑kernel macro‑texture cues with small‑kernel boundary details, and a Feature Additive Refinement Module (FARM) that globally refines fine‑grained anomaly representations. The authors also release the Copper Tube Defect Dataset (CTDD), a benchmark of 1,847 images with 4,898 annotated defect boxes. Experiments show CF‑YOLO outperforms baseline detectors such as YOLOv11 by 2.2% in mAP@50 and 3.9% in Precision while preserving real‑time speed.
This paper presents the IEEE International Conference on Multimedia and Expo (ICME) 2026 Grand Challenge on Cross-Scenario Defect Detection and Fine-Grained Severity Grading for High-Precision Manufacturing. The challenge is motivated by two key limitations of existing industrial defect inspection systems: (1) current deep learning-based methods often suffer significant performance degradation when deployed in unseen production scenarios, and (2) most benchmarks neglect severity-aware assessment, which is critical for risk control and yield optimization.
Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Diagnosing these failures requires instance-level feedback that answers where a defect occurs, what type it is, why it is defective, and its importance to overall image quality.
arXiv:2608.20713v1 Announce Type: new Abstract: Generative AI can now produce highly realistic images, yet current models still exhibit subtle but critical defects that undermine their reliability. W...
arXiv:2606. 01992v1 Announce Type: cross Abstract: Industrial anomaly detection has historically been a unimodal task.
arXiv:2607. 21155v1 Announce Type: cross Abstract: Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions.
arXiv:2507. 02288v2 Announce Type: replace-cross Abstract: Domain Generalization (DG) seeks to develop a versatile model capable of performing effectively on unseen target domains.
The paper introduces Anomaly‑LR, a defect‑grounded latent reasoning framework for industrial anomaly detection that builds a global understanding of an image and then refines anomaly‑relevant representations directly in visual latent space. It also presents IAD‑LR‑22K, a new instruction dataset with 22,228 image‑question pairs and detailed annotations. Experiments demonstrate that Anomaly‑LR outperforms comparable‑scale methods on multiple IAD benchmarks without needing external references or tools.
arXiv:2511. 01390v2 Announce Type: replace-cross Abstract: Fine-grained cross-modal alignment aims to establish precise local correspondences between vision and language, forming a cornerstone for visual question answering and related multimodal applications.
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. KI-VQA involves multiple sub-problems -referring expression understanding, visual grounding, object recognition, knowledge retrieval, and reasoning-yet existing benchmarks typically report only end-task accuracy, obscuring where failures arise.