The paper introduces TEEP‑RCNN, a two‑stage detector that augments Faster R‑CNN with a Feature Pyramid Network backbone and an enhanced Convolutional Block Attention Module (CBAM) featuring dropout in the channel attention MLP and batch‑norm in the spatial attention branch. Training employs a differential learning‑rate schedule with cosine‑annealing warm‑up, and inference uses Test‑Time Augmentation combined with Weighted Box Fusion to stabilize localization of elongated and boundary‑adjacent defects. On the NEU‑DET benchmark, TEEP‑RCNN attains 73.3 % mAP@50 and 37.9 % mAP@50‑95 in only ten epochs on a single GPU, matching or surpassing YOLOv11m while excelling on the rolled‑in‑scale defect category under the COCO metric.
By Kirtan Rajesh
arXiv:2608.20713v1 Announce Type: new
Abstract: Generative AI can now produce highly realistic images, yet current models still exhibit subtle but critical defects that undermine their reliability. W...
By Xiangfei Sheng, Weidong Zou, Tianjiao Gu, Zhichao Yang, Pengfei Chen, Leida Li
Vision Transformers underperform convolutional networks when training data is scarce, and distilling convolutional inductive biases from a CNN teacher is an effective remedy that leaves the deployed model unchanged. General-purpose feature distillation, however, transfers little in this setting.
The paper introduces TED (Text-Axis Evidence Decomposition), a post‑hoc scoring method that improves anomaly localization in CLIP‑based detectors without altering the backbone or prompts. TED evaluates whether ambiguous responses are better supported by defect patches or normal patches, thereby distinguishing true defects from visually complex normal regions. Experiments show that TED significantly enhances pixel‑level localization across frozen VLM backbones and adapted hosts, especially under hard‑false‑positive competition.
By JinYoung Kim, Geonho Kim, GiJeong Park, Geonu Lee, YoungJoon Yoo
Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Diagnosing these failures requires instance-level feedback that answers where a defect occurs, what type it is, why it is defective, and its importance to overall image quality.
arXiv:2609.23561v1 Announce Type: new
Abstract: Deploying efficient neural networks is essential in resource-constrained environments, yet compact models often sacrifice interpretability - a critical...
By Aleks Czufarow, Ihor Babin