arXiv Machine Learning By Peng Sun, Yi Yang, Antong Zhang, Chunxiao Li, Yanbo Wang, Dianbo Liu, xin chen, Kai Yu, Lu Chen, Tianfan Fu

Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training

Read the original on arXiv Machine Learning →

The paper introduces MASS, a hierarchical data selection method that first groups data using low‑dimensional manifold coordinates learned by a dense autoencoder, then applies a TopK sparse autoencoder for quality‑aware feature coverage within each group. This approach addresses the shortcomings of traditional diversity metrics that mix semantic, supervisory, and noise signals. Experiments on Vision Flan and LLaVA‑CoT demonstrate that MASS outperforms existing baselines across various budgets and can match or exceed full‑data training with only a small subset.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 25

Selective Inference for Deep Clustering in Latent Spaces

The paper introduces a selective inference framework tailored for deep clustering that uses a fixed pretrained encoder to map high‑dimensional data into a latent space before clustering. It addresses the complex selection bias arising from the nonlinear transformation and offers a computationally tractable method to perform valid statistical tests on cluster differences. Experiments on synthetic data show controlled Type I error and higher power compared to conservative baselines, while genomic case studies demonstrate the ability to uncover significant cluster differences while properly accounting for selection bias.

By Eina Mizui, Tomohiro Shiraishi, Shunichi Nishino, Ichiro Takeuchi
arXiv Computer Vision
Sep 3

KSG-Net: Key-Sparse and Global-Context Learning for Maritime 3D Ship Detection

KSG‑Net introduces a Key‑Sparse and Global‑Context learning framework for maritime 3D ship detection, addressing weak feature representation of small, sparse vessels and limited global modeling of large vessels. The network employs a Key Sparse Multi‑scale Aggregation module to select informative voxels and aggregate cross‑scale features, and a Global Context Aggregation module to capture long‑range geometric dependencies via scene‑level context modeling. Experiments on the Thames River vessel dataset and simulated data show that KSG‑Net outperforms existing methods in multi‑scale vessel detection and remains robust in complex maritime environments.

By Zhouyuan Huai, Meiqi Wan, Yan Yang, Minshi Chen, Xin Yuan, Wei Wang, Xiao Wang