arXiv Machine Learning

Analyzing Visual Aircraft Representations with Sparse Autoencoders

arXiv:2606. 15468v1 Announce Type: cross Abstract: Vision models can achieve strong performance on classification tasks, but the internal representations supporting their predictions are often difficult to interpret.

arXiv Computer Vision
Sep 3

KSG-Net: Key-Sparse and Global-Context Learning for Maritime 3D Ship Detection

KSG‑Net introduces a Key‑Sparse and Global‑Context learning framework for maritime 3D ship detection, addressing weak feature representation of small, sparse vessels and limited global modeling of large vessels. The network employs a Key Sparse Multi‑scale Aggregation module to select informative voxels and aggregate cross‑scale features, and a Global Context Aggregation module to capture long‑range geometric dependencies via scene‑level context modeling. Experiments on the Thames River vessel dataset and simulated data show that KSG‑Net outperforms existing methods in multi‑scale vessel detection and remains robust in complex maritime environments.

By Zhouyuan Huai, Meiqi Wan, Yan Yang, Minshi Chen, Xin Yuan, Wei Wang, Xiao Wang
arXiv Computer Vision
Sep 11

Beyond Benchmarks: Using VLMs to Reveal Systematic Classification Failures Under Real World Conditions

The paper investigates using Vision Language Models (VLMs) to accelerate verification and validation (V&V) of classification models by automatically detecting systematic errors. It introduces a VLM-based error slice detection (ESD) method that groups and labels errors, demonstrating its ability to identify perturbations in a non-military dataset and to cluster images by surroundings in a military context. The study highlights challenges such as underrepresentation of defence data in VLM training and limited contextual diversity, and suggests that while fully automated V&V is not yet feasible, VLMs could speed up the process in the future.

By Dieuwertje Alblas, Alma M. Liezenga, Jan Erik van Woerden, Fedor Taggenbrock, Dalia Aljawaheri, Klamer Schutte
arXiv AI
Jun 24

Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations

arXiv:2606. 24716v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence.

By Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Beg\"um Demir
arXiv AI
Jul 7

TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models

arXiv:2607. 04593v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large number of visual tokens fed to the language model.

By Riccardo Renzulli, Gabriele Spadaro, Shruthi Gowda, Alaa Eddine Mazouz, Van-Tam Nguyen