arXiv:2607. 08605v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept.
By Weiduo Liao, Yunqiao Yang, Ying Wei
KSG‑Net introduces a Key‑Sparse and Global‑Context learning framework for maritime 3D ship detection, addressing weak feature representation of small, sparse vessels and limited global modeling of large vessels. The network employs a Key Sparse Multi‑scale Aggregation module to select informative voxels and aggregate cross‑scale features, and a Global Context Aggregation module to capture long‑range geometric dependencies via scene‑level context modeling. Experiments on the Thames River vessel dataset and simulated data show that KSG‑Net outperforms existing methods in multi‑scale vessel detection and remains robust in complex maritime environments.
By Zhouyuan Huai, Meiqi Wan, Yan Yang, Minshi Chen, Xin Yuan, Wei Wang, Xiao Wang
arXiv:2608. 20134v1 Announce Type: cross Abstract: We present a novel view on feature evolution in Vision Transformers (ViTs) by visualizing the training process over two dimensions -- network depth (layer) and training time (epochs).
By Joonas J\"arve, Halil Ibrahim Aysel, Tarun Khajuria, Meelis Kull
arXiv:2608. 14922v1 Announce Type: cross Abstract: Mechanistic interpretability has recently expanded to Vision Transformers (ViTs), with Sparse Autoencoders (SAEs) increasingly used as post-hoc tools to decompose internal representations into sparse and more interpretable features.
By Philip H. Lee, Parth Padalkar
arXiv:2606. 06664v1 Announce Type: cross Abstract: Despite high accuracy, Vision Transformer (ViT) predictions can be driven by spurious cues, raising the need to understand their inner workings before safe deployment.
By Tang Li, Yanlin Chen, Mengmeng Ma, Xi Peng
The paper investigates using Vision Language Models (VLMs) to accelerate verification and validation (V&V) of classification models by automatically detecting systematic errors. It introduces a VLM-based error slice detection (ESD) method that groups and labels errors, demonstrating its ability to identify perturbations in a non-military dataset and to cluster images by surroundings in a military context. The study highlights challenges such as underrepresentation of defence data in VLM training and limited contextual diversity, and suggests that while fully automated V&V is not yet feasible, VLMs could speed up the process in the future.
By Dieuwertje Alblas, Alma M. Liezenga, Jan Erik van Woerden, Fedor Taggenbrock, Dalia Aljawaheri, Klamer Schutte