arXiv:2607. 12245v1 Announce Type: cross Abstract: Few-shot industrial defect detection remains difficult for standard supervised detectors, which achieve poor performance on boundary-dominated industrial defects.
By Jiaqi Kuang
CALIPER is a model‑free RGB‑D framework that performs fine‑grained recognition of visually similar industrial parts by combining support‑based appearance matching with metric size evidence. Each class is onboarded from a single turntable RGB‑D video and a few labeled real images, enabling 3D reconstruction for appearance support and depth‑aligned size profiling. At inference, a YOLOv8n‑seg model localizes parts, a frozen DINOv2 backbone with an episodically trained embedding head matches support, and margin‑conditioned metric fusion selectively uses size evidence for ambiguous cases, achieving high accuracy on 18 parts and robust enrollment of unseen screws without retraining.
By Alankrit Gupta, Chenxi Tao, Seung-Kyum Choi
arXiv:2604. 26633v2 Announce Type: replace-cross Abstract: Industrial surface defect inspection suffers from a fundamental data bottleneck: defects are rare, annotations require expert knowledge, and collecting balanced training sets is slow and costly.
By Paul Julius K\"uhn, Mika Pommeranz, Arjan Kuijper, Saptarshi Neil Sinha
The paper introduces TEEP‑RCNN, a two‑stage detector that augments Faster R‑CNN with a Feature Pyramid Network backbone and an enhanced Convolutional Block Attention Module (CBAM) featuring dropout in the channel attention MLP and batch‑norm in the spatial attention branch. Training employs a differential learning‑rate schedule with cosine‑annealing warm‑up, and inference uses Test‑Time Augmentation combined with Weighted Box Fusion to stabilize localization of elongated and boundary‑adjacent defects. On the NEU‑DET benchmark, TEEP‑RCNN attains 73.3 % mAP@50 and 37.9 % mAP@50‑95 in only ten epochs on a single GPU, matching or surpassing YOLOv11m while excelling on the rolled‑in‑scale defect category under the COCO metric.
By Kirtan Rajesh
arXiv:2603. 10834v3 Announce Type: replace-cross Abstract: Understanding how neural networks rely on visual cues offers a human-interpretable view of their internal decision processes.
By Pum Jun Kim, Seung-Ah Lee, Seongho Park, Dongyoon Han, Jaejun Yoo
The paper introduces the Aero-engine Blade Defect Detector (ABDD), an online adaptive detection framework that uses test-time adaptation to handle domain shifts in aero-engine blade inspection. ABDD employs a Dual-Alignment Strategy combining feature-statistics alignment with pseudo-box alignment to adapt both global visual style and local defect morphology, and incorporates an Uncertainty-aware Box Filtering mechanism to mitigate errors from noisy pseudo labels. A lightweight Sparse Dilated Mona module enables efficient parameter tuning while preventing source-domain forgetting, and the method is validated on CD-AeBD and HD-AeBD datasets, showing improved robustness and practical applicability on an industrial inspection platform.
By Zhaoyang Wang, Haiyong Chen, Dongying Li, Yining Wang, Huapeng Wu, Xinwei Lv, Atik Shahariar
Groundbench is a new benchmark that evaluates vision‑language models on multi‑resolution polygon grounding, using the same 1,500 image‑expression‑referent triples but targeting exact‑N polygons with five different vertex budgets. It audits both filled‑region intersection‑over‑union (IoU) and legal‑polygon completion, revealing that the best models achieve 88.2 box IoU and 97.1 accuracy at IoU ≥ 0.5, while direct polygon predictions lag at 57.7 and 69.2. The study shows performance is non‑monotonic across budgets, collapses at the densest budget due to legality failures, and highlights that false spatial cues hurt more than false colour cues, underscoring an operational geometry gap beyond latent boundary perception.
By Zhonghan Bian, Zhenran Wang, Jinsong Li, Zhangyang Qi
arXiv:2609.06232v1 Announce Type: cross
Abstract: Ground-truth defect masks in industrial inspection datasets are typically reserved for evaluation. This paper repurposes them as spatial supervision...
By Sajjad Rezvani Boroujeni, Muskan Saraf, Gnana Tulasi Makineni, Tom Bush, Hossein Abedi
arXiv:2606. 00844v1 Announce Type: cross Abstract: Bounding-box regression is a fundamental component of object detection, playing a critical role in precise object localization.
By Vinay Edula, Priyanka Bagade
Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Diagnosing these failures requires instance-level feedback that answers where a defect occurs, what type it is, why it is defective, and its importance to overall image quality.
CF-YOLO introduces a real‑time detection framework for camouflaged micro‑defects on industrial components, combining a Context‑Perception Aggregation Module (CPAM) that fuses large‑kernel macro‑texture cues with small‑kernel boundary details, and a Feature Additive Refinement Module (FARM) that globally refines fine‑grained anomaly representations. The authors also release the Copper Tube Defect Dataset (CTDD), a benchmark of 1,847 images with 4,898 annotated defect boxes. Experiments show CF‑YOLO outperforms baseline detectors such as YOLOv11 by 2.2% in mAP@50 and 3.9% in Precision while preserving real‑time speed.
By Xinda Yu, Kunxin Zheng, Chunan Yu, Qingbo Song, Hao Xiao, Ying Zang, Jie Liu
arXiv:2607. 12278v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) for computational pathology report striking zero-shot performance on whole-slide image (WSI) visual question answering (VQA) benchmarks.
By Wenhao Zhang, Zhongliang Zhou, John Kang, Sheng Li