arXiv AI

Rethinking Global Average Pooling: Your Classifier Is Secretly a Multi-Instance Learner

arXiv:2606. 14555v1 Announce Type: cross Abstract: Modern image classifiers widely adopt global average pooling (GAP) followed by a linear classification head.

arXiv AI
Sep 1

Background-Free Objectness Learning for Class-Agnostic Detection

Background-Free Objectness Learning (B-FOR) is a dense, class‑agnostic detection framework that learns objectness without treating unlabeled regions as background. It predicts multi‑scale object‑center and scale fields, using spatially structured soft targets to supervise only reliable annotated areas and introduces displacement‑aware scale fields to model object extent. Experiments on PASCAL VOC, MS‑COCO, and Open Images show B‑FOR improves recall by over +10 AR points compared to prior class‑agnostic baselines, with ablation studies confirming the importance of localized supervision and displacement‑aware scaling.

By Dania Batool, Liliana Lo Presti, Marco La Cascia, Filippo Vella
arXiv Machine Learning
Jul 3

Object-centric LeJEPA

arXiv:2607. 02404v1 Announce Type: cross Abstract: Image encoders trained with LeJEPA can deliver strong features for downstream tasks, but, like other image-level self-supervised methods, typically require large training datasets.

By Jakob Geusen, Ender Konukoglu
arXiv Machine Learning
Jul 14

Vertical Fusion: Condensing Internal Representations for Robust ViT Classification

arXiv:2607. 10391v1 Announce Type: cross Abstract: Despite exposing rich intermediate representations, Vision Transformers (ViTs) are almost exclusively utilized as black-box feature extractors, where only the last layer is considered for downstream tasks.

By Francesco Di Salvo, Shyam Nandan Rai, Hamed Damirchi, Ignacio Meza De la Jara, Sebastian Doerrich, Marco Lents, Christian Ledig
arXiv AI
Jul 7

Towards Generalizable Deepfake Image Detection with Vision Transformers

arXiv:2604. 17376v2 Announce Type: replace-cross Abstract: In today's day and age, we face a challenge in detecting deepfake images because of the fast evolution of modern generative models and the poor generalization capability of existing methods.

By Kaliki V Srinanda, M Manvith Prabhu, Hemanth K Mogilipalem, Jayavarapu S Abhinai, Vaibhav Santhosh, Aryan Herur, Deepu Vijayasenan
arXiv Computer Vision
Sep 18

Instance Segmentation and Fine-grained Classification for Urban Buildings with Adaptive Region Dividing and Spatially-Supervised Contrastive Learning

The paper introduces an adaptive region‑dividing strategy that projects a 3D point cloud onto a bird’s‑eye‑view plane to detect building regions, then back‑projects bounding boxes to create structure‑aligned training blocks for unified scene‑level evaluation. It also proposes a fine‑grained classification model using a point transformer classifier and a spatially‑supervised contrastive loss to improve inter‑class discriminability, addressing class imbalance with a weighted cross‑entropy. Experiments on UrbanBIS and STPLS3D datasets show the method outperforms state‑of‑the‑art approaches in both building instance segmentation and fine‑grained classification.

By Weiyuan Zhang, Qi Zhang, Hui Huang
arXiv Computer Vision
Aug 27

DeCO: Discriminative Evidence Composition for Fine-Grained Dataset Distillation

DeCO: Discriminative Evidence Composition for Fine-Grained Dataset Distillation introduces a method that compresses large training sets into compact synthetic sets while preserving fine-grained visual classification cues. It uses attention rollout from a pretrained TransFG teacher to locate informative patches, applies spatial diversification to avoid redundancy, and organizes these patches into class-wise evidence banks that are packed into grid-composed images. Experiments on CUB-200-2011, FGVC-Aircraft, and Stanford Cars demonstrate that DeCO outperforms existing coreset and dataset-distillation baselines across various images-per-class budgets.

By Chuixuan Fan, Guang Li, Shijie Wang, Dongzhan Zhou, Baoli Sun, Takahiro Ogawa, Miki Haseyama, Zhihui Wang