arXiv Machine Learning

A Realistic Protocol for Evaluation of Weakly Supervised Object Localization

arXiv:2404. 10034v3 Announce Type: replace-cross Abstract: Weakly Supervised Object Localization (WSOL) allows training deep learning models for classification and localization (LOC) using only global class-level labels.

arXiv AI
Sep 1

Background-Free Objectness Learning for Class-Agnostic Detection

Background-Free Objectness Learning (B-FOR) is a dense, class‑agnostic detection framework that learns objectness without treating unlabeled regions as background. It predicts multi‑scale object‑center and scale fields, using spatially structured soft targets to supervise only reliable annotated areas and introduces displacement‑aware scale fields to model object extent. Experiments on PASCAL VOC, MS‑COCO, and Open Images show B‑FOR improves recall by over +10 AR points compared to prior class‑agnostic baselines, with ablation studies confirming the importance of localized supervision and displacement‑aware scaling.

By Dania Batool, Liliana Lo Presti, Marco La Cascia, Filippo Vella
arXiv AI
Sep 10

WSPolypNet: Weakly Supervised Polyp Localization in Colonoscopy Videos

WSPolypNet is a weakly supervised framework that localizes polyps in colonoscopy videos using only video-level labels, avoiding costly frame-level annotations. It employs a 3D CNN to generate class activation maps, enhances them with a multi-view strategy, and refines the results with MedSAM2 segmentation. The method achieves higher CorLoc scores—up to 47.80% at IoU 0.3—and a recall of 94.51%, especially improving detection of small polyps.

By Giseong Hwang, Minjae Jo, Yeonghyeon Park, Kyeonghun Kim, Seoyeon Han, Donghoon Han, Haneul Kim, Yului Jeong, Insung Hwang, Pa Hong, Ken Ying-Kai Liao, Nam-Joon Kim
arXiv Computer Vision
Sep 18

Distance to Class Prototypes: Active Learning for Object Detection

The paper introduces a new active learning signal for object detection that relies on a supervised contrastive term added to the training objective. This term shapes an embedding space where distance reflects class membership, allowing an unlabeled detection to be scored by its distance from the predicted category’s region weighted by confidence—all from a single forward pass of one network. Experiments on PASCAL VOC and MS‑COCO show that this criterion outperforms the standard posterior and remains competitive with ensemble‑based methods while incurring only a modest 8.3% increase in parameters.

By Licheng Zhang, Zheng Gong
arXiv Computer Vision
Sep 4

Solving the Needle-in-a-Haystack Problem in Mammography Vision-Language Model with Differentiable Subset Sampling

The paper introduces TopKSigLIP, a vision‑language model tailored for mammography that tackles two key challenges: high‑resolution imaging and homogeneous radiology reports. It replaces standard CLIP training with a TopK‑Patch module that selects sparse high‑resolution patches likely to contain lesions, and a Sup‑sigmoid loss that uses soft labels from structured data instead of contrastive loss. TopKSigLIP outperforms existing open‑source mammography and general medical VLMs on zero‑shot tasks such as density assessment, BI‑RADS classification, finding subtyping, and cancer prediction, while also providing better lesion localization than Grad‑CAM.

By Young Seok Jeon, Beatrice Brown-Mulry, Rohan Satya Isaac, Anjana Dissanayaka, Theo Dapamede, Mohammadreza Chavoshi, Judy Gichoya, Hari Trivedi
arXiv Computer Vision
Sep 24

OD3: Optimization-free Dataset Distillation for Object Detection

OD3 introduces an optimization‑free dataset distillation framework tailored for object detection. The method first iteratively places object instances in synthesized images, then screens candidates with a pre‑trained observer model to discard low‑confidence objects. Applied to MS COCO and PASCAL VOC, OD3 achieves compression ratios from 0.25% to 5% and surpasses previous detection‑focused distillation methods by over 14% on COCO mAP50 at a 1.0% compression ratio.

By Salwa K. Al Khatib, Ahmed ElHagry, Shitong Shao, Zhiqiang Shen