Background-Free Objectness Learning (B-FOR) is a dense, class‑agnostic detection framework that learns objectness without treating unlabeled regions as background. It predicts multi‑scale object‑center and scale fields, using spatially structured soft targets to supervise only reliable annotated areas and introduces displacement‑aware scale fields to model object extent. Experiments on PASCAL VOC, MS‑COCO, and Open Images show B‑FOR improves recall by over +10 AR points compared to prior class‑agnostic baselines, with ablation studies confirming the importance of localized supervision and displacement‑aware scaling.
By Dania Batool, Liliana Lo Presti, Marco La Cascia, Filippo Vella
arXiv:2501.12632v3 Announce Type: replace-cross
Abstract: Weakly supervised object localization (WSOL) models can predict both the object class and the spatial regions corresponding to the object, wi...
By Shakeeb Murtaza, Soufiane Belharbi, Alexis Guichemerre, Marco Pedersoli, Eric Granger
arXiv:2610.01778v1 Announce Type: new
Abstract: Reliable evaluation of image forgery localization (IFL) requires assessing models under diverse distribution changes, yet existing benchmarks often cov...
By Baoke Dou, Ziye Wang, Hao Wang, Guoqing Cai, Wende Tan, Chenyang Si, Liucheng Guo, Yueming Lyu
WSPolypNet is a weakly supervised framework that localizes polyps in colonoscopy videos using only video-level labels, avoiding costly frame-level annotations. It employs a 3D CNN to generate class activation maps, enhances them with a multi-view strategy, and refines the results with MedSAM2 segmentation. The method achieves higher CorLoc scores—up to 47.80% at IoU 0.3—and a recall of 94.51%, especially improving detection of small polyps.
By Giseong Hwang, Minjae Jo, Yeonghyeon Park, Kyeonghun Kim, Seoyeon Han, Donghoon Han, Haneul Kim, Yului Jeong, Insung Hwang, Pa Hong, Ken Ying-Kai Liao, Nam-Joon Kim
arXiv:2606. 00844v1 Announce Type: cross Abstract: Bounding-box regression is a fundamental component of object detection, playing a critical role in precise object localization.
By Vinay Edula, Priyanka Bagade
arXiv:2609.25500v1 Announce Type: new
Abstract: Training data quantity and quality greatly affect object detection model performance, regardless of model architecture. When using object detection mod...
By Lonny Lundsten, Kevin Barnard, Dave Caress
The paper introduces a new active learning signal for object detection that relies on a supervised contrastive term added to the training objective. This term shapes an embedding space where distance reflects class membership, allowing an unlabeled detection to be scored by its distance from the predicted category’s region weighted by confidence—all from a single forward pass of one network. Experiments on PASCAL VOC and MS‑COCO show that this criterion outperforms the standard posterior and remains competitive with ensemble‑based methods while incurring only a modest 8.3% increase in parameters.
By Licheng Zhang, Zheng Gong
The paper introduces TopKSigLIP, a vision‑language model tailored for mammography that tackles two key challenges: high‑resolution imaging and homogeneous radiology reports. It replaces standard CLIP training with a TopK‑Patch module that selects sparse high‑resolution patches likely to contain lesions, and a Sup‑sigmoid loss that uses soft labels from structured data instead of contrastive loss. TopKSigLIP outperforms existing open‑source mammography and general medical VLMs on zero‑shot tasks such as density assessment, BI‑RADS classification, finding subtyping, and cancer prediction, while also providing better lesion localization than Grad‑CAM.
By Young Seok Jeon, Beatrice Brown-Mulry, Rohan Satya Isaac, Anjana Dissanayaka, Theo Dapamede, Mohammadreza Chavoshi, Judy Gichoya, Hari Trivedi
arXiv:2609.36882v1 Announce Type: new
Abstract: Localizing AI-edited regions is essential for interpretable forensic analysis, but remains challenging due to subtle and spatially distributed artifact...
By Junhee Lee, Donghyeon Jeon, Taeoh Kim, Beomyoung Kim, MyeongAh Cho
arXiv:2607. 10783v1 Announce Type: cross Abstract: Whole-slide images (WSIs) provide rich tissue-level and cellular-level information, but storing and transmitting high-magnification pathology data is resource-intensive.
By Dung Minh Do, Nhat-Thanh Huynh, Duc Minh Huynh, Doanh C. Bui, Khang Nguyen
arXiv:2509.22404v2 Announce Type: replace
Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; howe...
By Yiwei Li, Yikang Liu, Jiaqi Guo, Lin Zhao, Zheyuan Zhang, Xiao Chen, Boris Mailhe, Ankush Mukherjee, Terrence Chen, Shanhui Sun
OD3 introduces an optimization‑free dataset distillation framework tailored for object detection. The method first iteratively places object instances in synthesized images, then screens candidates with a pre‑trained observer model to discard low‑confidence objects. Applied to MS COCO and PASCAL VOC, OD3 achieves compression ratios from 0.25% to 5% and surpasses previous detection‑focused distillation methods by over 14% on COCO mAP50 at a 1.0% compression ratio.
By Salwa K. Al Khatib, Ahmed ElHagry, Shitong Shao, Zhiqiang Shen