arXiv Computation and Language

Vague2Detect: Handling Ambiguous Prompts in Knowledge-Based Open-World Detection

Vague2Detect is a hybrid pipeline that improves object detection for ambiguous prompts by combining a fine‑tuned Sentence‑BERT to retrieve candidates from a structured household knowledge base, YOLO‑World to verify their presence in images, and a GPT‑3.5‑turbo fallback to generate new candidate descriptions when prompts fall outside the knowledge base. On a benchmark of household scenes, Vague2Detect raises the vague prompt success rate from 32% (YOLO‑World alone) to 61% with high precision, and up to 85% when the GPT fallback is used.

arXiv Computer Vision
Sep 3

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker

The paper introduces OpenRef, a benchmark for Referring Expression Comprehension (REC) designed for open‑world scenarios. OpenRef expands beyond simple settings by including diverse visual domains, variable target counts (multi‑target and none‑target), and a rich vocabulary with proper nouns, polysemous words, and ordinal terms. It also proposes new evaluation metrics—F1 for grounding accuracy and N3R for negative expression rejection—and presents a training‑free Multi‑task Consistency Checker (MCC) that improves model performance with a single click.

By Zongjian Wu, Lei Zhang
arXiv Machine Learning
Jun 15

Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes

arXiv:2602. 00593v4 Announce Type: replace-cross Abstract: Despite progress on general tasks, vision-language models (VLMs) still struggle with challenges that demand both fine-grained visual grounding and external knowledge, a synergy overlooked by existing benchmarks that evaluate these abilities in isolation.

By Yifan Jiang, Cong Zhang, Bofei Zhang, Qiaofeng Zheng, Yifan Yang, Bingzhang Wang, Yew-Soon Ong
arXiv Computer Vision
Aug 28

CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection

CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection introduces a unified inference-time framework that addresses semantic ambiguity and over-suppression in multimodal OWOD systems. It comprises Cross-Modal Joint Confidence Calibration, Uncertainty-Guided Universal Objectness Enhancement, and Dynamic Outlier Suppression via Confidence Margin. Experiments on the Real-World Detection benchmark with the OWL‑ViT L/14 backbone show CODE achieving 21.7 U‑mAP and 40.8 K‑mAP, surpassing prior state‑of‑the‑art results by 2.6 and 2.3 points respectively.

By Hao Xu, Zhaoning Shi, Hehe Jin, Bo Ma
Hugging Face Trending Papers
Jul 23

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. KI-VQA involves multiple sub-problems -referring expression understanding, visual grounding, object recognition, knowledge retrieval, and reasoning-yet existing benchmarks typically report only end-task accuracy, obscuring where failures arise.