Hugging Face Trending Papers

Vague2Detect: Handling Ambiguous Prompts in Knowledge-Based Open-World Detection

arXiv Computation and Language
Sep 10

Vague2Detect: Handling Ambiguous Prompts in Knowledge-Based Open-World Detection

Vague2Detect is a hybrid pipeline that improves object detection for ambiguous prompts by combining a fine‑tuned Sentence‑BERT to retrieve candidates from a structured household knowledge base, YOLO‑World to verify their presence in images, and a GPT‑3.5‑turbo fallback to generate new candidate descriptions when prompts fall outside the knowledge base. On a benchmark of household scenes, Vague2Detect raises the vague prompt success rate from 32% (YOLO‑World alone) to 61% with high precision, and up to 85% when the GPT fallback is used.

By Ibrohimjon Muminov (Dongguk University, Seoul, South Korea), Jihie Kim (Dongguk University, Seoul, South Korea)
arXiv Computer Vision
Sep 3

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker

The paper introduces OpenRef, a benchmark for Referring Expression Comprehension (REC) designed for open‑world scenarios. OpenRef expands beyond simple settings by including diverse visual domains, variable target counts (multi‑target and none‑target), and a rich vocabulary with proper nouns, polysemous words, and ordinal terms. It also proposes new evaluation metrics—F1 for grounding accuracy and N3R for negative expression rejection—and presents a training‑free Multi‑task Consistency Checker (MCC) that improves model performance with a single click.

By Zongjian Wu, Lei Zhang
arXiv Computer Vision
Aug 28

CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection

CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection introduces a unified inference-time framework that addresses semantic ambiguity and over-suppression in multimodal OWOD systems. It comprises Cross-Modal Joint Confidence Calibration, Uncertainty-Guided Universal Objectness Enhancement, and Dynamic Outlier Suppression via Confidence Margin. Experiments on the Real-World Detection benchmark with the OWL‑ViT L/14 backbone show CODE achieving 21.7 U‑mAP and 40.8 K‑mAP, surpassing prior state‑of‑the‑art results by 2.6 and 2.3 points respectively.

By Hao Xu, Zhaoning Shi, Hehe Jin, Bo Ma
arXiv Computer Vision
Sep 21

Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty

The paper demonstrates that object detection benchmarks suffer from incomplete annotations, with re-annotation of COCO, Pascal VOC, Cityscapes, and KITTI revealing up to a 60% increase in detected objects, especially small, occluded, or densely packed instances. The authors propose a scalable annotation pipeline that uses multiple annotators per object to capture uncertainty and improve recall, and they introduce two new large-scale benchmarks: an uncertainty-aware detection benchmark and a label error detection benchmark based on real errors. Their findings show that benchmark performance is highly sensitive to annotation quality, yet model rankings remain largely unchanged, highlighting the need for uncertainty-aware evaluation to better reflect real-world ambiguity.

By Sarina Penquitt, Jonathan Klees, Antonia van Betteray, Parssa Jashnieh, Peter Stehr, Matthias Rottmann, Lars Schmarje
Hugging Face Trending Papers
Jul 1

GKDT: General Keypoint Detection Transformer

With the emergence of various pre-trained vision and language models, computer vision is shifting from narrow-domain to open-domain recognition. The construction of a more powerful yet general keypoint detection (GKD) model to support diverse tasks has become increasingly important in the field.