Real-world detectors must often interpret functional or ambiguous prompts, yet conventional models such as YOLO remain restricted to fixed class lists. Even open-vocabulary models like YOLO-World freq...
arXiv:2604. 17376v2 Announce Type: replace-cross Abstract: In today's day and age, we face a challenge in detecting deepfake images because of the fast evolution of modern generative models and the poor generalization capability of existing methods.
By Kaliki V Srinanda, M Manvith Prabhu, Hemanth K Mogilipalem, Jayavarapu S Abhinai, Vaibhav Santhosh, Aryan Herur, Deepu Vijayasenan
arXiv:2605.15088v2 Announce Type: replace
Abstract: Detecting 3D keypoints is a long-standing challenge in computer vision. Most detectors end with a heuristic post-processing step that is not learne...
By Batuhan Arda Bekar, Can Sar{\i}, H\"useyin Can G\"ulkan, Bar{\i}\c{s} \"Ozcan
arXiv:2607. 24745v1 Announce Type: cross Abstract: Key Information Extraction (KIE) is vital for many document applications, but creating training datasets is traditionally a time-consuming manual process.
By Siddartha Reddy, Harikrishnan P M, Goutham Vignesh, Varun V, Vishal Vaddina
Vague2Detect is a hybrid pipeline that improves object detection for ambiguous prompts by combining a fine‑tuned Sentence‑BERT to retrieve candidates from a structured household knowledge base, YOLO‑World to verify their presence in images, and a GPT‑3.5‑turbo fallback to generate new candidate descriptions when prompts fall outside the knowledge base. On a benchmark of household scenes, Vague2Detect raises the vague prompt success rate from 32% (YOLO‑World alone) to 61% with high precision, and up to 85% when the GPT fallback is used.
By Ibrohimjon Muminov (Dongguk University, Seoul, South Korea), Jihie Kim (Dongguk University, Seoul, South Korea)
arXiv:2607.02486v2 Announce Type: replace
Abstract: Descriptor-free visual localization eliminates high-dimensional descriptor storage, preserves scene privacy, and simplifies map maintenance, yet it...
By Yejun Zhang, Xinjue Wang, Zihan Wang, Esa Rahtu, Juho Kannala
Crane is a CLIP‑based framework for zero‑shot anomaly detection that enhances dense localization by adapting the vision encoder with a correlation‑based attention module and conditioning learnable prompts on global image context. It further fuses anomaly‑relevant patch features into the global representation for more sensitive image‑level detection, and a variant called Crane+ leverages DINOv2 spatial correlations for stronger pixel‑level performance. Across seven industrial benchmarks, Crane raises mean image‑level AP by 4.5% and Crane+ boosts mean pixel‑level AUPRO by 9.0%.
By Alireza Salehi, Mohammadreza Salehi, Reshad Hosseini, Cees G. M. Snoek, Makoto Yamada, Mohammad Sabokrou
arXiv:2607. 09876v1 Announce Type: cross Abstract: Automatically retrieving videos from large camera-trap datasets remains challenging.
By Valentin Gabeff, Baptiste Maquignaz, Jennifer Shan, Sepideh Mamooler, Gencer Sumbul, Blair Costelloe, Devis Tuia, Alexander Mathis
arXiv:2609.35490v2 Announce Type: replace
Abstract: Generalist multitasking vision models aim to unify multiple vision tasks within a single framework, enabling more efficient and versatile learning....
By Mohammad Mahdi, Nedyalko Prisadnikov, Yuqian Fu, Carmelo Scribano, Danda Pani Paudel, Luc Van Gool
arXiv:2604. 02327v2 Announce Type: replace-cross Abstract: Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification, and segmentation.
By Jona Ruthardt, Manu Gaur, Deva Ramanan, Makarand Tapaswi, Yuki M. Asano
arXiv:2609.23431v1 Announce Type: new
Abstract: Human-object interaction (HOI) detection requires grounding an interacting human-object pair and recognizing the verb that links them, often under seve...
By Junwen Chen, Keiji Yanai
arXiv:2607.09086v2 Announce Type: replace
Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...
By Jie Zhu, Ivy Zhang, Minchul Kim, Xiaoming Liu