Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes. PROB improves unknown discovery by modeling class-agnostic probabilistic objectness in the decoder-query space.
Vague2Detect is a hybrid pipeline that improves object detection for ambiguous prompts by combining a fine‑tuned Sentence‑BERT to retrieve candidates from a structured household knowledge base, YOLO‑World to verify their presence in images, and a GPT‑3.5‑turbo fallback to generate new candidate descriptions when prompts fall outside the knowledge base. On a benchmark of household scenes, Vague2Detect raises the vague prompt success rate from 32% (YOLO‑World alone) to 61% with high precision, and up to 85% when the GPT fallback is used.
By Ibrohimjon Muminov (Dongguk University, Seoul, South Korea), Jihie Kim (Dongguk University, Seoul, South Korea)
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. KI-VQA involves multiple sub-problems -referring expression understanding, visual grounding, object recognition, knowledge retrieval, and reasoning-yet existing benchmarks typically report only end-task accuracy, obscuring where failures arise.
The paper introduces Structured Prior Knowledge (SPK), a framework that extracts and organizes latent priors from pretrained object detectors to improve out-of-distribution (OoD) detection. SPK uses in-distribution data and hallucination-inducing samples to elicit part-level semantic concepts, then combines these with geometric and contextual priors into a compact five-dimensional representation. Experiments across various detector architectures and OoD benchmarks show that SPK achieves state-of-the-art performance, demonstrating that pretrained detectors encode richer latent knowledge than previously exploited.
By Changshun Wu, Weicheng He, Xiaowei Huang, Saddek Bensalem
The paper demonstrates that object detection benchmarks suffer from incomplete annotations, with re-annotation of COCO, Pascal VOC, Cityscapes, and KITTI revealing up to a 60% increase in detected objects, especially small, occluded, or densely packed instances. The authors propose a scalable annotation pipeline that uses multiple annotators per object to capture uncertainty and improve recall, and they introduce two new large-scale benchmarks: an uncertainty-aware detection benchmark and a label error detection benchmark based on real errors. Their findings show that benchmark performance is highly sensitive to annotation quality, yet model rankings remain largely unchanged, highlighting the need for uncertainty-aware evaluation to better reflect real-world ambiguity.
By Sarina Penquitt, Jonathan Klees, Antonia van Betteray, Parssa Jashnieh, Peter Stehr, Matthias Rottmann, Lars Schmarje
arXiv:2609.27076v1 Announce Type: new
Abstract: Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queries without dependence on predefined perce...
By Linus Nwankwo, Muslim Alaran, Christian Rauch, Stanley Chukwuebuka Obilikpa, Elmar Rueckert
arXiv:2607. 21155v1 Announce Type: cross Abstract: Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions.
By Hanseok Oh, Parishad BehnamGhader, Benno Krojer, Hyunji Lee, Paul Liang, Siva Reddy, Verna Dankers
arXiv:2609.23431v1 Announce Type: new
Abstract: Human-object interaction (HOI) detection requires grounding an interacting human-object pair and recognizing the verb that links them, often under seve...
By Junwen Chen, Keiji Yanai
arXiv:2607. 04548v1 Announce Type: cross Abstract: Novel category discovery aims to identify unseen classes from unlabeled data by transferring knowledge from labeled categories, but most existing methods perform discovery in opaque latent feature spaces.
By Ifrat Ikhtear Uddin, Yang Zhou, KC Santosh, Longwei Wang
Real-world detectors must often interpret functional or ambiguous prompts, yet conventional models such as YOLO remain restricted to fixed class lists. Even open-vocabulary models like YOLO-World freq...
arXiv:2609.12552v1 Announce Type: new
Abstract: Open-vocabulary detection accepts any class list at inference, and promptable segmentation returns regions without class names: the taxonomy has left t...
By Ma\"elic Neau
arXiv:2607. 01759v1 Announce Type: cross Abstract: Open-vocabulary object detection aims to localize and classify objects beyond the fixed set of categories seen dur ing training.
By Jae-Ryung Hong, Ho-Joong Kim, Seong-Whan Lee