arXiv:2604. 06435v2 Announce Type: replace-cross Abstract: Visual Anomaly Detection (VAD) is a critical task for many applications including industrial inspection and healthcare.
By Manuel Barusco, Francesco Borsatti, David Petrovic, Davide Dalle Pezze, Gian Antonio Susto
arXiv:2504.10214v2 Announce Type: replace
Abstract: Pretrained model-based incremental object detection (PTMIOD) leverages the rich detection priors of pretrained detectors to learn new categories in...
By Songze Li, Qixing Xu, Tonghua Su, Xu-Yao Zhang, Zhongjie Wang, Yunzhe Li
Segmentation-based multi-object tracking (MOT) with foundation video models such as SAM2 offers strong localization quality, yet remains fragile in crowded, real-world scenes. In detector-prompted SAM...
arXiv:2510.21175v2 Announce Type: replace
Abstract: Pre-trained vision-language models (VLMs), such as CLIP, have demonstrated remarkable zero-shot generalization, enabling deployment in a wide range...
By Yujin Jo, Taesup Kim
LiAM‑SAM is a lifecycle‑aware memory framework designed to improve segmentation‑based multi‑object tracking (MOT) with the SAM2 foundation video model. It addresses three common failure modes—faulty track initiation, memory drift during close interactions, and unreliable re‑identification after occlusion—by introducing contrastive track initiation, motion‑ and geometry‑grounded memory correction, and adaptive context memory. The system achieves state‑of‑the‑art HOTA and IDF1 scores, with ablations showing significant gains in association metrics and a 96% reduction in identity switches.
By Gr\'egoire Francisco, Alessandro D'Amico, Samuele Costantini, Gianpiero Francesca, Lorenzo Garattoni
arXiv:2607. 17157v1 Announce Type: cross Abstract: Multi-object tracking (MOT) aims to localize multiple objects in videos while preserving their identities over time.
By Yanrong Qin, Xiaoyan Cao, Yao Yao
The paper investigates how rectifying supermarket product images using homography estimation and the Hough transform can improve deep learning-based object detection. It evaluates the impact of angle variation and object density on detection accuracy, highlighting both benefits and limitations of image rectification. The authors advocate for a new dataset to further study these effects.
By Mayank Sah, Jimson Mathew
CF-YOLO introduces a real‑time detection framework for camouflaged micro‑defects on industrial components, combining a Context‑Perception Aggregation Module (CPAM) that fuses large‑kernel macro‑texture cues with small‑kernel boundary details, and a Feature Additive Refinement Module (FARM) that globally refines fine‑grained anomaly representations. The authors also release the Copper Tube Defect Dataset (CTDD), a benchmark of 1,847 images with 4,898 annotated defect boxes. Experiments show CF‑YOLO outperforms baseline detectors such as YOLOv11 by 2.2% in mAP@50 and 3.9% in Precision while preserving real‑time speed.
By Xinda Yu, Kunxin Zheng, Chunan Yu, Qingbo Song, Hao Xiao, Ying Zang, Jie Liu
arXiv:2510. 06596v2 Announce Type: replace-cross Abstract: The performance of machine learning models depends heavily on training data.
By Ayush Zenith, Arnold Zumbrun, Neel Raut, Jing Lin
CALIPER is a model‑free RGB‑D framework that performs fine‑grained recognition of visually similar industrial parts by combining support‑based appearance matching with metric size evidence. Each class is onboarded from a single turntable RGB‑D video and a few labeled real images, enabling 3D reconstruction for appearance support and depth‑aligned size profiling. At inference, a YOLOv8n‑seg model localizes parts, a frozen DINOv2 backbone with an episodically trained embedding head matches support, and margin‑conditioned metric fusion selectively uses size evidence for ambiguous cases, achieving high accuracy on 18 parts and robust enrollment of unseen screws without retraining.
By Alankrit Gupta, Chenxi Tao, Seung-Kyum Choi
arXiv:2607. 09785v1 Announce Type: cross Abstract: Traditionally, continual learning has assumed access to labeled data, yet many real-world applications -- such as lifelong robotics -- require models to adapt continuously from unlabeled streams.
By Sergi Masip, Alicja Dobrzeniecka, Jonathan Swinnen, Joachim Collin, Bart{\l}omiej Twardowski, Szymon {\L}ukasik, Tinne Tuytelaars
arXiv:2601. 22012v3 Announce Type: replace Abstract: Catastrophic forgetting in continual learning is often measured at the performance or last-layer representation level, overlooking the underlying mechanisms.
By Sergi Masip, Gido M. van de Ven, Javier Ferrando, Tinne Tuytelaars