Background-Free Objectness Learning (B-FOR) is a dense, class‑agnostic detection framework that learns objectness without treating unlabeled regions as background. It predicts multi‑scale object‑center and scale fields, using spatially structured soft targets to supervise only reliable annotated areas and introduces displacement‑aware scale fields to model object extent. Experiments on PASCAL VOC, MS‑COCO, and Open Images show B‑FOR improves recall by over +10 AR points compared to prior class‑agnostic baselines, with ablation studies confirming the importance of localized supervision and displacement‑aware scaling.
By Dania Batool, Liliana Lo Presti, Marco La Cascia, Filippo Vella
arXiv:2607. 02404v1 Announce Type: cross Abstract: Image encoders trained with LeJEPA can deliver strong features for downstream tasks, but, like other image-level self-supervised methods, typically require large training datasets.
By Jakob Geusen, Ender Konukoglu
arXiv:2607. 10391v1 Announce Type: cross Abstract: Despite exposing rich intermediate representations, Vision Transformers (ViTs) are almost exclusively utilized as black-box feature extractors, where only the last layer is considered for downstream tasks.
By Francesco Di Salvo, Shyam Nandan Rai, Hamed Damirchi, Ignacio Meza De la Jara, Sebastian Doerrich, Marco Lents, Christian Ledig
arXiv:2604. 17376v2 Announce Type: replace-cross Abstract: In today's day and age, we face a challenge in detecting deepfake images because of the fast evolution of modern generative models and the poor generalization capability of existing methods.
By Kaliki V Srinanda, M Manvith Prabhu, Hemanth K Mogilipalem, Jayavarapu S Abhinai, Vaibhav Santhosh, Aryan Herur, Deepu Vijayasenan
Vision-language models (VLMs) such as CLIP enable zero-shot classification by comparing image features with text prompts in a shared embedding space. A fundamental property underlying this capability is the global comparability of logits across arbitrary candidate classes.
arXiv:2609.38603v1 Announce Type: new
Abstract: While earth observation models have advanced substantially, they still lack interpretability. While concept-bottleneck models provide interpretability...
By Rishabh Mondal, Nipun Batra, Utkarsh Mall
The paper introduces an adaptive region‑dividing strategy that projects a 3D point cloud onto a bird’s‑eye‑view plane to detect building regions, then back‑projects bounding boxes to create structure‑aligned training blocks for unified scene‑level evaluation. It also proposes a fine‑grained classification model using a point transformer classifier and a spatially‑supervised contrastive loss to improve inter‑class discriminability, addressing class imbalance with a weighted cross‑entropy. Experiments on UrbanBIS and STPLS3D datasets show the method outperforms state‑of‑the‑art approaches in both building instance segmentation and fine‑grained classification.
By Weiyuan Zhang, Qi Zhang, Hui Huang
DeCO: Discriminative Evidence Composition for Fine-Grained Dataset Distillation introduces a method that compresses large training sets into compact synthetic sets while preserving fine-grained visual classification cues. It uses attention rollout from a pretrained TransFG teacher to locate informative patches, applies spatial diversification to avoid redundancy, and organizes these patches into class-wise evidence banks that are packed into grid-composed images. Experiments on CUB-200-2011, FGVC-Aircraft, and Stanford Cars demonstrate that DeCO outperforms existing coreset and dataset-distillation baselines across various images-per-class budgets.
By Chuixuan Fan, Guang Li, Shijie Wang, Dongzhan Zhou, Baoli Sun, Takahiro Ogawa, Miki Haseyama, Zhihui Wang
arXiv:2607. 05978v1 Announce Type: cross Abstract: Multimodal large language models can emit localized predictions, bounding boxes for objects and temporal windows for video and audio events, but they hallucinate these regions prolifically.
By Daniel Shalam, Emanuel Ben Baruch, Avi Ben Cohen, Tal Remez
arXiv:2607.09086v2 Announce Type: replace
Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...
By Jie Zhu, Ivy Zhang, Minchul Kim, Xiaoming Liu
arXiv:2609.13706v1 Announce Type: new
Abstract: Co-salient object detection (Co-SOD) requires a model to find foreground regions that are salient in individual images and supported by the image group...
By Yuan Xiang, Matteo Rossi, Yingzhou Chen
arXiv:2608.20548v1 Announce Type: cross
Abstract: Disaster damage is spatial: buildings rarely fail in isolation. Yet using spatial context for damage classification remains surprisingly underexplore...
By Fuad Hasan, Chul Min Yeum