OPUS: A Simple yet Effective Unified Framework for Open-Vocabulary Detection
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2609.23431v1 Announce Type: new Abstract: Human-object interaction (HOI) detection requires grounding an interacting human-object pair and recognizing the verb that links them, often under seve...
VIVAS is a new Vision‑Language Model pre‑training framework that addresses the lack of fine‑grained visual perception in existing VLMs. It introduces a unified token space and a dense‑structural‑semantic vision tokenizer that expands the textual vocabulary with visual tokens, enabling vision‑language unified autoregressive supervision over both visual details and linguistic content. Trained on 12.4 T tokens, VIVAS achieves state‑of‑the‑art results on 7 tasks and 39 multimodal benchmarks.
arXiv:2609.27076v1 Announce Type: new Abstract: Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queries without dependence on predefined perce...
ProCAP introduces a probabilistic cross-attentive prompt learning framework for vision-language models like CLIP, enabling improved cross-modal interaction without updating the backbone. It jointly learns visual and textual prompt tokens, linking them via stacked bidirectional multi-head cross-attention to refine each branch across prompt depth. The method incorporates Gaussian parameterization of prompt tokens, lightweight KL and L2 regularization, and a compact symmetric InfoNCE head to align image features with class-level text representations, achieving strong few-shot base-to-novel performance and competitive transfer results across multiple datasets and benchmarks.
arXiv:2608.28707v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Visual Question Answering (VQA), yet they continue to struggle with quest...
Efficient Unified Multimodal Understanding (EUMU) is the winning solution for the MUMU Track of the 8th LSVOS Challenge, addressing multi‑concept image tagging, open‑vocabulary object detection, and image captioning with a single efficient model. It leverages a shared pretrained multimodal backbone and lightweight heads, while applying task‑aware inference refinement that uses detection cues to improve captioning, caption cues to recover missed detections, and image statistics to refine tagging. With 239.169 M parameters, 23.947 GFLOPs, and 4.5 GB peak memory, EUMU achieves a challenge score of 17.3409 and is publicly available on GitHub.