The paper introduces the concept of information density to explain category bias in visual object detection. It finds a strong negative correlation between a category’s information density and its detection accuracy, showing that instance count alone does not account for bias. By incorporating information density into three advanced loss functions, the authors demonstrate reduced model bias and improved overall performance on Pascal VOC, COCO‑LT, and LVIS datasets.
By Ziwei Zhao, Yanxi Lu, Yuwei Hu, Shiyang Su, Mingxuan Wang, Chenyue Zhou, Jiayi Chen, Hehan Li, Xiaoshuai Hao, Andi Zhang, Yanbiao Ma
Balancing performance trade-offs on long-tailed data distributions remains a long-standing challenge in visual recognition. Existing methods mainly improve tail classes through re-balancing, represent...
arXiv:2609.37331v1 Announce Type: new
Abstract: Balancing performance trade-offs on long-tailed data distributions remains a long-standing challenge in visual recognition. Existing methods mainly imp...
By Shenghan Chen, Yiming Liu, Zhipeng Deng, Haolin Wang, Jiale Zhou, Zhijian Wu, Xiankai Lu, Yafei Ou, Yefeng Zheng
arXiv:2510. 06596v2 Announce Type: replace-cross Abstract: The performance of machine learning models depends heavily on training data.
By Ayush Zenith, Arnold Zumbrun, Neel Raut, Jing Lin
arXiv:2609.14654v1 Announce Type: new
Abstract: Convolutional neural networks (CNNs) are typically evaluated using held-out classification accuracy, an approach that presupposes predictions are based...
By Abhilekha Dalal, Michael Okonoda, Eder Martinez, Lior Shamir
arXiv:2605. 07821v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is crucial for ensuring the reliability of deep learning models.
By Boyang Dai, Chaoqi Chen, Yizhou Yu
arXiv:2608. 00716v1 Announce Type: cross Abstract: Robust detection of generated images is critical to counter the misuse of generative models.
By Jun Nie, Yonggang Zhang, Tongliang Liu, Yiu-ming Cheung, Bo Han, Xinmei Tian
arXiv:2605. 00273v2 Announce Type: replace-cross Abstract: Text-to-image diffusion models achieve impressive visual fidelity, yet they remain unreliable in multi-object generation.
By Yujin Jeong, Arnas Uselis, Iro Laina, Seong Joon Oh, Anna Rohrbach
The paper demonstrates that object detection benchmarks suffer from incomplete annotations, with re-annotation of COCO, Pascal VOC, Cityscapes, and KITTI revealing up to a 60% increase in detected objects, especially small, occluded, or densely packed instances. The authors propose a scalable annotation pipeline that uses multiple annotators per object to capture uncertainty and improve recall, and they introduce two new large-scale benchmarks: an uncertainty-aware detection benchmark and a label error detection benchmark based on real errors. Their findings show that benchmark performance is highly sensitive to annotation quality, yet model rankings remain largely unchanged, highlighting the need for uncertainty-aware evaluation to better reflect real-world ambiguity.
By Sarina Penquitt, Jonathan Klees, Antonia van Betteray, Parssa Jashnieh, Peter Stehr, Matthias Rottmann, Lars Schmarje
Background-Free Objectness Learning (B-FOR) is a dense, class‑agnostic detection framework that learns objectness without treating unlabeled regions as background. It predicts multi‑scale object‑center and scale fields, using spatially structured soft targets to supervise only reliable annotated areas and introduces displacement‑aware scale fields to model object extent. Experiments on PASCAL VOC, MS‑COCO, and Open Images show B‑FOR improves recall by over +10 AR points compared to prior class‑agnostic baselines, with ablation studies confirming the importance of localized supervision and displacement‑aware scaling.
By Dania Batool, Liliana Lo Presti, Marco La Cascia, Filippo Vella
arXiv:2606. 00082v1 Announce Type: cross Abstract: Explainability of deep learning algorithms is critical for computer-vision applications with high-stake decisions.
By Cl\'ement B\'enard, Manon Arfib, Christophe Labreuche, Victor Qu\'etu
arXiv:2607. 02404v1 Announce Type: cross Abstract: Image encoders trained with LeJEPA can deliver strong features for downstream tasks, but, like other image-level self-supervised methods, typically require large training datasets.
By Jakob Geusen, Ender Konukoglu