arXiv AI

Morphology-Aware Sample Assignment: Overcoming IoU Insensitivity for Surface Defect Detection

arXiv:2606. 13723v1 Announce Type: cross Abstract: Intersection-over-Union (IoU), as a pivotal metric for evaluating the spatial alignment between candidate proposals and ground-truth annotations, directly determines the quality of positive sample sets and the training efficacy of visual detection models.

arXiv Computer Vision
Sep 17

CALIPER: Metric-Grounded Model-Free Recognition of Visually Similar Industrial Parts

CALIPER is a model‑free RGB‑D framework that performs fine‑grained recognition of visually similar industrial parts by combining support‑based appearance matching with metric size evidence. Each class is onboarded from a single turntable RGB‑D video and a few labeled real images, enabling 3D reconstruction for appearance support and depth‑aligned size profiling. At inference, a YOLOv8n‑seg model localizes parts, a frozen DINOv2 backbone with an episodically trained embedding head matches support, and margin‑conditioned metric fusion selectively uses size evidence for ambiguous cases, achieving high accuracy on 18 parts and robust enrollment of unseen screws without retraining.

By Alankrit Gupta, Chenxi Tao, Seung-Kyum Choi
arXiv AI
Jul 23

SynSur: An end-to-end generative pipeline for synthetic industrial surface defect generation and detection

arXiv:2604. 26633v2 Announce Type: replace-cross Abstract: Industrial surface defect inspection suffers from a fundamental data bottleneck: defects are rare, annotations require expert knowledge, and collecting balanced training sets is slow and costly.

By Paul Julius K\"uhn, Mika Pommeranz, Arjan Kuijper, Saptarshi Neil Sinha
arXiv Computer Vision
Sep 24

TEEP-RCNN: Texture-Enhanced Edge-aware Perception for Steel Surface Defect Detection via Improved Convolutional Block Attention in Faster R-CNN

The paper introduces TEEP‑RCNN, a two‑stage detector that augments Faster R‑CNN with a Feature Pyramid Network backbone and an enhanced Convolutional Block Attention Module (CBAM) featuring dropout in the channel attention MLP and batch‑norm in the spatial attention branch. Training employs a differential learning‑rate schedule with cosine‑annealing warm‑up, and inference uses Test‑Time Augmentation combined with Weighted Box Fusion to stabilize localization of elongated and boundary‑adjacent defects. On the NEU‑DET benchmark, TEEP‑RCNN attains 73.3 % mAP@50 and 37.9 % mAP@50‑95 in only ten epochs on a single GPU, matching or surpassing YOLOv11m while excelling on the rolled‑in‑scale defect category under the COCO metric.

By Kirtan Rajesh
arXiv Computer Vision
2d ago

Robust Online Aero-Engine Blade Defect Detection via Dual-Alignment Test-Time Adaptation

The paper introduces the Aero-engine Blade Defect Detector (ABDD), an online adaptive detection framework that uses test-time adaptation to handle domain shifts in aero-engine blade inspection. ABDD employs a Dual-Alignment Strategy combining feature-statistics alignment with pseudo-box alignment to adapt both global visual style and local defect morphology, and incorporates an Uncertainty-aware Box Filtering mechanism to mitigate errors from noisy pseudo labels. A lightweight Sparse Dilated Mona module enables efficient parameter tuning while preventing source-domain forgetting, and the method is validated on CD-AeBD and HD-AeBD datasets, showing improved robustness and practical applicability on an industrial inspection platform.

By Zhaoyang Wang, Haiyong Chen, Dongying Li, Yining Wang, Huapeng Wu, Xinwei Lv, Atik Shahariar
arXiv Computer Vision
Sep 24

Groundbench: Multi-Resolution Polygon Grounding Exposes the Geometry Gap in Vision-Language Models

Groundbench is a new benchmark that evaluates vision‑language models on multi‑resolution polygon grounding, using the same 1,500 image‑expression‑referent triples but targeting exact‑N polygons with five different vertex budgets. It audits both filled‑region intersection‑over‑union (IoU) and legal‑polygon completion, revealing that the best models achieve 88.2 box IoU and 97.1 accuracy at IoU ≥ 0.5, while direct polygon predictions lag at 57.7 and 69.2. The study shows performance is non‑monotonic across budgets, collapses at the densest budget due to legality failures, and highlights that false spatial cues hurt more than false colour cues, underscoring an operational geometry gap beyond latent boundary perception.

By Zhonghan Bian, Zhenran Wang, Jinsong Li, Zhangyang Qi
arXiv Machine Learning
Sep 10

Spatial Attention Supervision for Defect Localization: Exploiting Ground-Truth Masks as Training Signal in Diffusion-Augmented Defect Detection

arXiv:2609.06232v1 Announce Type: cross Abstract: Ground-truth defect masks in industrial inspection datasets are typically reserved for evaluation. This paper repurposes them as spatial supervision...

By Sajjad Rezvani Boroujeni, Muskan Saraf, Gnana Tulasi Makineni, Tom Bush, Hossein Abedi
Hugging Face Trending Papers
Jun 4

Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Diagnosing these failures requires instance-level feedback that answers where a defect occurs, what type it is, why it is defective, and its importance to overall image quality.

arXiv Computer Vision
Aug 31

CF-YOLO: Context-Aware Feature Refinement for Camouflaged Industrial Micro-Defect Detection

CF-YOLO introduces a real‑time detection framework for camouflaged micro‑defects on industrial components, combining a Context‑Perception Aggregation Module (CPAM) that fuses large‑kernel macro‑texture cues with small‑kernel boundary details, and a Feature Additive Refinement Module (FARM) that globally refines fine‑grained anomaly representations. The authors also release the Copper Tube Defect Dataset (CTDD), a benchmark of 1,847 images with 4,898 annotated defect boxes. Experiments show CF‑YOLO outperforms baseline detectors such as YOLOv11 by 2.2% in mAP@50 and 3.9% in Precision while preserving real‑time speed.

By Xinda Yu, Kunxin Zheng, Chunan Yu, Qingbo Song, Hao Xiao, Ying Zang, Jie Liu