arXiv:2608. 20810v1 Announce Type: cross Abstract: Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its production, so the output must remain retrievable by the queries it was meant to serve.
By Guangyuan Dong, Chuang Liu, Yangchen Zeng, Haoyu Wang, Xiaoyang Yu, Pinlong Zhao, Yuchao Hou, Ziwei Li, Zheng Lin
Background-Free Objectness Learning (B-FOR) is a dense, class‑agnostic detection framework that learns objectness without treating unlabeled regions as background. It predicts multi‑scale object‑center and scale fields, using spatially structured soft targets to supervise only reliable annotated areas and introduces displacement‑aware scale fields to model object extent. Experiments on PASCAL VOC, MS‑COCO, and Open Images show B‑FOR improves recall by over +10 AR points compared to prior class‑agnostic baselines, with ablation studies confirming the importance of localized supervision and displacement‑aware scaling.
By Dania Batool, Liliana Lo Presti, Marco La Cascia, Filippo Vella
The paper introduces CERES, a closed‑loop multimodal indexing framework that addresses semantic collapse in multimodal generation by building a three‑level semantic pyramid and using scale‑routed cross‑attention to generate images that remain retrievable by their original queries. CERES employs a co‑occurrence‑aware router, a lightweight U‑Net generator, and a soft‑Jaccard coverage objective to ensure generated images cover the intended concepts, verified by re‑indexing with a frozen vision‑language model and an external DINOv2 probe. Experiments on four pansharpening benchmarks show state‑of‑the‑art performance, especially under extreme scale variation, and significant improvements in concept‑query retrieval and image‑text ranking metrics.
By Guangyuan Dong, Chuang Liu, Haoyu Wang, Yangchen Zeng, Jiaqi Zhang, Li Jiuxing, Xiaoyang Yu, Pinlong Zhao, Yuchao Hou, Ziwei Li, Zheng Lin, Alexander Lim Han Yang, Yusen Wu
The paper introduces Structured Prior Knowledge (SPK), a framework that extracts and organizes latent priors from pretrained object detectors to improve out-of-distribution (OoD) detection. SPK uses in-distribution data and hallucination-inducing samples to elicit part-level semantic concepts, then combines these with geometric and contextual priors into a compact five-dimensional representation. Experiments across various detector architectures and OoD benchmarks show that SPK achieves state-of-the-art performance, demonstrating that pretrained detectors encode richer latent knowledge than previously exploited.
By Changshun Wu, Weicheng He, Xiaowei Huang, Saddek Bensalem
Metric questions about video require vision-language models to use supplied real-world references to convert visual measurements into physical units. Yet we find that current models use this scale inf...
The paper introduces EquiSD, a label‑free training method that exploits scale equivariance to improve metric grounding in vision‑language models. By projecting model predictions onto a scale‑equivariant family and fine‑tuning on the resulting targets, EquiSD boosts a 3B model’s median response slope from 0.66 to 0.94 and raises mean relative accuracy by 9.2 points across simulated scales, with positive transfer to real QuantiPhy videos.
By Kaizhen Tan, Yang Feng, Heqing Du, Siru Tao, Xin Xu, Hanzhe Hong
The paper demonstrates that object detection benchmarks suffer from incomplete annotations, with re-annotation of COCO, Pascal VOC, Cityscapes, and KITTI revealing up to a 60% increase in detected objects, especially small, occluded, or densely packed instances. The authors propose a scalable annotation pipeline that uses multiple annotators per object to capture uncertainty and improve recall, and they introduce two new large-scale benchmarks: an uncertainty-aware detection benchmark and a label error detection benchmark based on real errors. Their findings show that benchmark performance is highly sensitive to annotation quality, yet model rankings remain largely unchanged, highlighting the need for uncertainty-aware evaluation to better reflect real-world ambiguity.
By Sarina Penquitt, Jonathan Klees, Antonia van Betteray, Parssa Jashnieh, Peter Stehr, Matthias Rottmann, Lars Schmarje
The paper introduces PromptCCZSL, a framework that enables vision‑language models to continually learn new attributes, objects, and their unique compositions while avoiding forgetting. It uses a frozen VLM backbone with prompt‑based techniques, recency‑weighted multi‑teacher distillation, and several loss functions (CAL, OPL, IDL) to maintain prior knowledge and promote diverse, distinct embeddings. Experiments on UT‑Zappos and C‑GQA show significant performance gains over existing VLM‑based and non‑VLM baselines, establishing a new benchmark for continual compositional zero‑shot learning.
By Sauda Maryam, Sara Nadeem, Faisal Qureshi, Mohsen Ali
arXiv:2603. 09493v2 Announce Type: replace-cross Abstract: The adaptation of large-scale vision-language models (VLMs) to downstream tasks with limited labeled data remains a significant challenge.
By Enming Zhang, Jiayang Li, Yanlong Wang, Yanru Wu, Zhenyu Liu, Yang Li
Text-to-image (T2I) diffusion models have achieved striking progress but still struggle to synthesize rare concepts involving unusual attribute-object pairings, often resulting in concept omission or semantic drift where a dominant entity overwhelms the generation. Tracing these failures to a lack of compositional balance during the denoising trajectory, we propose RADIANCE, a training-free framework that treats inference as a closed-loop feedback process.
arXiv:2607. 00371v1 Announce Type: cross Abstract: Visual AutoRegressive modeling (VAR) has pioneered a coarse-to-fine multi-scale autoregressive generative paradigm, demonstrating strong capabilities in image generation.
By Nuoyan Zhou, Zhijun Tu, Lei Yu, Kun Cheng, Jie Hu, Nannan Wang, Xinghao Chen
arXiv:2504.10214v2 Announce Type: replace
Abstract: Pretrained model-based incremental object detection (PTMIOD) leverages the rich detection priors of pretrained detectors to learn new categories in...
By Songze Li, Qixing Xu, Tonghua Su, Xu-Yao Zhang, Zhongjie Wang, Yunzhe Li