arXiv Machine Learning

Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training

The paper introduces MASS, a hierarchical data selection method that first groups data using low‑dimensional manifold coordinates learned by a dense autoencoder, then applies a TopK sparse autoencoder for quality‑aware feature coverage within each group. This approach addresses the shortcomings of traditional diversity metrics that mix semantic, supervisory, and noise signals. Experiments on Vision Flan and LLaVA‑CoT demonstrate that MASS outperforms existing baselines across various budgets and can match or exceed full‑data training with only a small subset.

arXiv Machine Learning
Sep 25

Selective Inference for Deep Clustering in Latent Spaces

The paper introduces a selective inference framework tailored for deep clustering that uses a fixed pretrained encoder to map high‑dimensional data into a latent space before clustering. It addresses the complex selection bias arising from the nonlinear transformation and offers a computationally tractable method to perform valid statistical tests on cluster differences. Experiments on synthetic data show controlled Type I error and higher power compared to conservative baselines, while genomic case studies demonstrate the ability to uncover significant cluster differences while properly accounting for selection bias.

By Eina Mizui, Tomohiro Shiraishi, Shunichi Nishino, Ichiro Takeuchi
arXiv Computer Vision
Sep 3

KSG-Net: Key-Sparse and Global-Context Learning for Maritime 3D Ship Detection

KSG‑Net introduces a Key‑Sparse and Global‑Context learning framework for maritime 3D ship detection, addressing weak feature representation of small, sparse vessels and limited global modeling of large vessels. The network employs a Key Sparse Multi‑scale Aggregation module to select informative voxels and aggregate cross‑scale features, and a Global Context Aggregation module to capture long‑range geometric dependencies via scene‑level context modeling. Experiments on the Thames River vessel dataset and simulated data show that KSG‑Net outperforms existing methods in multi‑scale vessel detection and remains robust in complex maritime environments.

By Zhouyuan Huai, Meiqi Wan, Yan Yang, Minshi Chen, Xin Yuan, Wei Wang, Xiao Wang
arXiv Machine Learning
Aug 31

A Deeper Analysis of Block-Sparse Featurizers

The paper investigates the block-sparse featurizer (BSF), a model that uses small subspaces as atomic units instead of single directions, aiming to capture features on low-dimensional manifolds common in vision. It identifies that BSF still exhibits classic sparse autoencoder failure modes such as feature splitting and composition. The authors propose architectural improvements, notably a Tournament Top‑K selection rule, which markedly reduces feature splitting, and they extend the block concept to a crosscoder framework.

By Alexandru-Iulius Jerpelea, Amith Ananthram
arXiv Machine Learning
Aug 31

EXPOSE: Explainable and Domain-Robust Embeddings from Pathology Vision Foundation Models using Sparse Autoencoders

EXPOSE is a framework that applies Sparse Autoencoders to Vision Foundation Model embeddings in computational pathology, aiming to separate biological signals from domain‑specific noise. By training a sparse representation of VFM features and using a linear classifier to flag domain‑specific latent dimensions, the method masks these components before downstream relapse prediction, avoiding the need to retrain the backbone model. Experiments on a large prostate cancer dataset demonstrate that removing domain‑specific features improves cross‑domain performance and raises the Domain Robustness Index (DoRI).

By Anja Witte, Maximilian Lennartz, Jan Baumbach, Guido Sauter, Stefan Bonn, Patrick Fuhlert, Marina Zimmermann
Hugging Face Trending Papers
Aug 13

CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport

While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe computational bottlenecks during inference. Existing token pruning methods primarily rely on diversity-based selection, discarding similar tokens to maximize dispersion.