Real-world detectors must often interpret functional or ambiguous prompts, yet conventional models such as YOLO remain restricted to fixed class lists. Even open-vocabulary models like YOLO-World freq...
Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes. PROB improves unknown discovery by modeling class-agnostic probabilistic objectness in the decoder-query space.
arXiv:2607. 23981v1 Announce Type: cross Abstract: Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes.
By Weijun Tian, Rui Liu
The paper introduces OpenRef, a benchmark for Referring Expression Comprehension (REC) designed for open‑world scenarios. OpenRef expands beyond simple settings by including diverse visual domains, variable target counts (multi‑target and none‑target), and a rich vocabulary with proper nouns, polysemous words, and ordinal terms. It also proposes new evaluation metrics—F1 for grounding accuracy and N3R for negative expression rejection—and presents a training‑free Multi‑task Consistency Checker (MCC) that improves model performance with a single click.
By Zongjian Wu, Lei Zhang
arXiv:2602. 00593v4 Announce Type: replace-cross Abstract: Despite progress on general tasks, vision-language models (VLMs) still struggle with challenges that demand both fine-grained visual grounding and external knowledge, a synergy overlooked by existing benchmarks that evaluate these abilities in isolation.
By Yifan Jiang, Cong Zhang, Bofei Zhang, Qiaofeng Zheng, Yifan Yang, Bingzhang Wang, Yew-Soon Ong
arXiv:2608.30247v1 Announce Type: new
Abstract: Recent unified open-vocabulary detection (OVD) supports heterogeneous prompts, including text queries, visual exemplars, and their combinations, but of...
By Xiaoyan Wei, Zhimin Yao, Ruilin Yang, Wei Zhang, Yong Dai, Yi Zhang, Wei Ge
arXiv:2609.27076v1 Announce Type: new
Abstract: Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queries without dependence on predefined perce...
By Linus Nwankwo, Muslim Alaran, Christian Rauch, Stanley Chukwuebuka Obilikpa, Elmar Rueckert
arXiv:2609.24026v1 Announce Type: new
Abstract: In this paper, we investigate the limitations of fixed, hand-crafted connectors in hierarchical semantic representations for open-vocabulary object det...
By Yeong-Jin Kim, Ho-Joong Kim, Seong-Whan Lee
CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection introduces a unified inference-time framework that addresses semantic ambiguity and over-suppression in multimodal OWOD systems. It comprises Cross-Modal Joint Confidence Calibration, Uncertainty-Guided Universal Objectness Enhancement, and Dynamic Outlier Suppression via Confidence Margin. Experiments on the Real-World Detection benchmark with the OWL‑ViT L/14 backbone show CODE achieving 21.7 U‑mAP and 40.8 K‑mAP, surpassing prior state‑of‑the‑art results by 2.6 and 2.3 points respectively.
By Hao Xu, Zhaoning Shi, Hehe Jin, Bo Ma
arXiv:2608.28707v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Visual Question Answering (VQA), yet they continue to struggle with quest...
By Anoop Senthil
arXiv:2609.23431v1 Announce Type: new
Abstract: Human-object interaction (HOI) detection requires grounding an interacting human-object pair and recognizing the verb that links them, often under seve...
By Junwen Chen, Keiji Yanai
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. KI-VQA involves multiple sub-problems -referring expression understanding, visual grounding, object recognition, knowledge retrieval, and reasoning-yet existing benchmarks typically report only end-task accuracy, obscuring where failures arise.