arXiv:2606. 30344v1 Announce Type: cross Abstract: Visual classifiers can achieve high matched-distribution accuracy while relying on low-level cues that fail under conflict or suppression.
By Chanho Park, Woochan Lee, Janyeong Oh, Geongho Gong, Minshu Kim, Yeachan Kwak, Seongim Choi
arXiv:2603. 10834v3 Announce Type: replace-cross Abstract: Understanding how neural networks rely on visual cues offers a human-interpretable view of their internal decision processes.
By Pum Jun Kim, Seung-Ah Lee, Seongho Park, Dongyoon Han, Jaejun Yoo
The paper introduces TEEP‑RCNN, a two‑stage detector that augments Faster R‑CNN with a Feature Pyramid Network backbone and an enhanced Convolutional Block Attention Module (CBAM) featuring dropout in the channel attention MLP and batch‑norm in the spatial attention branch. Training employs a differential learning‑rate schedule with cosine‑annealing warm‑up, and inference uses Test‑Time Augmentation combined with Weighted Box Fusion to stabilize localization of elongated and boundary‑adjacent defects. On the NEU‑DET benchmark, TEEP‑RCNN attains 73.3 % mAP@50 and 37.9 % mAP@50‑95 in only ten epochs on a single GPU, matching or surpassing YOLOv11m while excelling on the rolled‑in‑scale defect category under the COCO metric.
By Kirtan Rajesh
The paper investigates how knowledge distillation from event cameras to RGB images can alter the inductive biases of convolutional neural networks. By transferring learning from the event domain, the authors find that models gain color invariance, a shape bias, and improved robustness to high‑frequency noise, largely due to reduced reliance on texture and increased emphasis on edge‑based object shape. These changes are evidenced by early‑layer processing differences and a spectral trade‑off between robustness to missing high‑frequency content and vulnerability to its contamination or geometric disruption.
By Soshun Kihara, Shunsuke Yasuki, Masato Taki
arXiv:2511.11286v4 Announce Type: replace-cross
Abstract: Out-of-domain (OOD) robustness is challenging to achieve in real-world computer vision, especially in unsupervised domain adaptation scenario...
By Ruoqi Wang, Haitao Wang, Shaojie Guo, Qiong Luo
arXiv:2608. 03557v1 Announce Type: cross Abstract: Tabular-to-image methods that convert tabular data into visual representations have emerged as a novel paradigm for leveraging the high performance of deep learning models.
By Malena Loza, Felipe Grijalva, Eva Milara, Luis Bote-Curiel, Francisco J. Lara-Abelenda, David Chushig-Muzo
The paper introduces a Focal Log-Frequency Loss (f-loss) to counteract the spectral imbalance in pixel-space flow matching, where low frequencies dominate training. By balancing learning signals across frequencies and combining early frequency-domain supervision with later pixel-space refinement, the method accelerates convergence by up to 40% and improves FID and perceptual fidelity across multiple model scales. It requires no architectural changes and can replace existing flow matching losses as a drop‑in solution.
By Lucas Degeorge, Paul Couairon, Arijit Ghosh, Alexei A. Efros, David Picard, Vicky Kalogeiton
arXiv:2608. 05424v1 Announce Type: cross Abstract: Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals.
By Vladan Stojni\'c, Ryan Ramos, Giorgos Kordopatis-Zilos, Noa Garcia, Giorgos Tolias
arXiv:2608. 09091v1 Announce Type: cross Abstract: Transfer learning is particularly useful in settings with limited training data, and within image classification it is common to transfer learn upon massive datasets like ImageNet , CIFAR-100, or COCO .
By Jing Ning, James D. Braza
arXiv:2608.29167v1 Announce Type: new
Abstract: Robustness of segmentation models is commonly assessed through input-domain perturbations, while dependence on frequency content within learned feature...
By Subhash Kashyap
arXiv:2606. 23825v1 Announce Type: cross Abstract: Efficient small object detection is bottlenecked by the inherent feature scarcity of tiny targets, which is further aggravated by operations of spatial-domain detectors that indiscriminately discard critical high-frequency details.
By Yuhan Rui, Shihan Qiao, Yibin Lou, Mingxi Yu, Yutong Wan, Yanqiao Chen, Dongsheng Hou, Zhen Cao, Athena Zhuoming Zhong, Qi Hao
arXiv:2204. 14224v3 Announce Type: replace-cross Abstract: The automated analysis of heterogeneous natural textures is frequently hindered by physical damage and data loss, presenting a significant challenge to computer vision.
By Galymzhan Abdimanap, Kairat Bostanbekov, Abdelrahman Abdallah, Anel Alimova, Darkhan Kurmangaliyev, Daniyar Nurseitov, Tatyana Dedova, Larissa Balakay, Serik Nurakynov