arXiv Machine Learning

Training-Free Human-in-the-Loop Anomaly Detection via Memory Bank Correction

The paper introduces a training‑free, human‑in‑the‑loop anomaly detection framework that allows a domain expert to correct a PatchCore detector by editing its memory bank, without retraining or using gradients. Using only ten golden samples, operator corrections close a median 66% of the performance gap to a fully trained bank, improving 12 of 15 MVTec AD categories while harming none. The approach is evaluated with a rigorous held‑out protocol and shows that passive and active querying yield statistically indistinguishable gains, with a defect‑memory extension failing decisively.

arXiv Computer Vision
Aug 25

When More References Hurt: Contamination-Aware DINOv2 Memory Banks for Few-Shot Steel Defect Detection

The paper investigates how to improve patch‑memory anomaly detectors for steel defect detection when additional industrial images may contain unseen defects. By filtering out the most suspicious 20% of patches from a contaminated reference bank and merging the remaining patches with a clean seed bank, the authors reduce contamination from 9.46% to 2.59% and achieve higher AUPRC scores compared to naive expansion or random removal. The method demonstrates that reference purity is a critical design factor and that unverified images can be beneficial only after explicit filtering.

By Hannaneh Kalantari, Javad Khoramdel
arXiv Computer Vision
Aug 24

When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference

arXiv:2608.21098v1 Announce Type: new Abstract: Fusing prior knowledge with data-driven learning is attractive where data is scarce, yet no controlled account says when it helps, is redundant, or har...

By Ahmad AlMughrabi, Albert Clop, Benjamin Busam, Ricardo Marques, Petia Radeva
arXiv Computer Vision
Sep 4

SafeRestore: Detector-Relative Risk Certificates for Selective Industrial Image Restoration

SafeRestore introduces a framework for certifying when an industrial image restoration should be automatically returned to a detector or require human review. It ranks five restoration candidates using action‑specific fitted scores, selects a threshold gate on tuning data, and evaluates the gate on a separate certification sample with two one‑sided exact binomial bounds—one for evidence‑loss incidents and one for excess‑activation incidents. In a retrospective study of 4,591 Carinthia‑S images, the protocol demonstrates auditable risk‑coverage behavior, with varying pass rates across different policies and morphologies.

By Shaoliang Yang, Jun Wang
arXiv Machine Learning
Sep 25

When Identical Rows Disagree: From Benchmark Identifiability to Replication-Robust Anomaly Detection

The paper investigates how repeated rows in released datasets—often treated as i.i.d. samples—introduce a hidden measurement layer that affects anomaly detection. It shows that identical rows can cap evaluation performance, make AUROC sensitive to replication, and bias detectors toward multiplicity size. The authors audit 690 OddBench datasets, find significant train-test overlap and label conflicts, and propose SCOUT, a support‑count orthogonalized detector that separates replication‑invariant evidence from exposure‑aware counts, achieving comparable or better AUROC while controlling false‑positive rates.

By Jie Deng
arXiv AI
Sep 4

Differentiable Interval Bottlenecks for Interpretable Anomaly Detection in Numerical Data

DIFFINT is a reconstruction‑based anomaly detector that uses a differentiable autoencoder with a latent bottleneck composed of soft, axis‑aligned interval memberships. Each latent unit represents a human‑readable hyper‑rectangle in feature space, allowing the model to encode how strongly an instance falls inside each interval and to compute reconstruction error as the anomaly score. The method provides a certified lower bound on reconstruction error for points outside all active intervals, a suppression mechanism for sparse abnormalities, and a closed‑form, label‑free importance ranking for each (unit, feature) pair, achieving top performance on 48 ADBench benchmarks against 22 baselines.

By Lamine Diop, Marc Plantevit