Vague2Detect is a hybrid pipeline that improves object detection for ambiguous prompts by combining a fine‑tuned Sentence‑BERT to retrieve candidates from a structured household knowledge base, YOLO‑World to verify their presence in images, and a GPT‑3.5‑turbo fallback to generate new candidate descriptions when prompts fall outside the knowledge base. On a benchmark of household scenes, Vague2Detect raises the vague prompt success rate from 32% (YOLO‑World alone) to 61% with high precision, and up to 85% when the GPT fallback is used.
By Ibrohimjon Muminov (Dongguk University, Seoul, South Korea), Jihie Kim (Dongguk University, Seoul, South Korea)
Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes. PROB improves unknown discovery by modeling class-agnostic probabilistic objectness in the decoder-query space.
The paper introduces OpenRef, a benchmark for Referring Expression Comprehension (REC) designed for open‑world scenarios. OpenRef expands beyond simple settings by including diverse visual domains, variable target counts (multi‑target and none‑target), and a rich vocabulary with proper nouns, polysemous words, and ordinal terms. It also proposes new evaluation metrics—F1 for grounding accuracy and N3R for negative expression rejection—and presents a training‑free Multi‑task Consistency Checker (MCC) that improves model performance with a single click.
By Zongjian Wu, Lei Zhang
arXiv:2607. 23981v1 Announce Type: cross Abstract: Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes.
By Weijun Tian, Rui Liu
arXiv:2608.30247v1 Announce Type: new
Abstract: Recent unified open-vocabulary detection (OVD) supports heterogeneous prompts, including text queries, visual exemplars, and their combinations, but of...
By Xiaoyan Wei, Zhimin Yao, Ruilin Yang, Wei Zhang, Yong Dai, Yi Zhang, Wei Ge
arXiv:2609.24026v1 Announce Type: new
Abstract: In this paper, we investigate the limitations of fixed, hand-crafted connectors in hierarchical semantic representations for open-vocabulary object det...
By Yeong-Jin Kim, Ho-Joong Kim, Seong-Whan Lee
CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection introduces a unified inference-time framework that addresses semantic ambiguity and over-suppression in multimodal OWOD systems. It comprises Cross-Modal Joint Confidence Calibration, Uncertainty-Guided Universal Objectness Enhancement, and Dynamic Outlier Suppression via Confidence Margin. Experiments on the Real-World Detection benchmark with the OWL‑ViT L/14 backbone show CODE achieving 21.7 U‑mAP and 40.8 K‑mAP, surpassing prior state‑of‑the‑art results by 2.6 and 2.3 points respectively.
By Hao Xu, Zhaoning Shi, Hehe Jin, Bo Ma
arXiv:2509.24192v2 Announce Type: replace
Abstract: Vision-language models (VLMs) have advanced multimodal perception, demonstrated by open-vocabulary object detection with simple language queries. S...
By Sojung An, Kwanyong Park, Yong Jae Lee, Donghyun Kim
arXiv:2608.28707v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Visual Question Answering (VQA), yet they continue to struggle with quest...
By Anoop Senthil
The paper demonstrates that object detection benchmarks suffer from incomplete annotations, with re-annotation of COCO, Pascal VOC, Cityscapes, and KITTI revealing up to a 60% increase in detected objects, especially small, occluded, or densely packed instances. The authors propose a scalable annotation pipeline that uses multiple annotators per object to capture uncertainty and improve recall, and they introduce two new large-scale benchmarks: an uncertainty-aware detection benchmark and a label error detection benchmark based on real errors. Their findings show that benchmark performance is highly sensitive to annotation quality, yet model rankings remain largely unchanged, highlighting the need for uncertainty-aware evaluation to better reflect real-world ambiguity.
By Sarina Penquitt, Jonathan Klees, Antonia van Betteray, Parssa Jashnieh, Peter Stehr, Matthias Rottmann, Lars Schmarje
With the emergence of various pre-trained vision and language models, computer vision is shifting from narrow-domain to open-domain recognition. The construction of a more powerful yet general keypoint detection (GKD) model to support diverse tasks has become increasingly important in the field.
arXiv:2609.23431v1 Announce Type: new
Abstract: Human-object interaction (HOI) detection requires grounding an interacting human-object pair and recognizing the verb that links them, often under seve...
By Junwen Chen, Keiji Yanai