arXiv:2606. 14592v1 Announce Type: cross Abstract: Clustering is widely used for exploratory analysis and scientific discovery, driving insights from market segmentation to biological data analysis, but its outputs can be difficult to interpret, audit, and reproduce as modern datasets become increasingly large and complex.
By Claire M. He, Genevera I. Allen
arXiv:2512. 17678v2 Announce Type: replace-cross Abstract: Selecting compact and informative gene subsets from single-cell transcriptomic data is essential for biomarker discovery, improving interpretability, and cost-effective profiling.
By Daphn\'e Chopard, Jorge da Silva Gon\c{c}alves, Irene Cannistraci, Thomas M. Sutter, Julia E. Vogt
arXiv:2509. 25289v4 Announce Type: replace-cross Abstract: Identifying an effective clustering algorithm for a given dataset remains a fundamental unsupervised learning issue.
By Mohammadreza Bakhtyari, Bogdan Mazoure, Renato Cordeiro de Amorim, Guillaume Rabusseau, Vladimir Makarenkov
The paper introduces Inverted Contrastive Learning for Unsupervised Feature Selection (ICLFS), a method that treats each feature as a sample by inverting the data matrix and applies a contrastive learning framework to learn consistent representations across masked positive views and a shuffled negative view. Feature saliency is derived from the magnitude of projector‑space embeddings, and a Laplacian‑Gated Ranking Correction step refines the ranking by reducing local redundancy. Experiments on 12 benchmark datasets show that ICLFS achieves the best clustering accuracy on 10 datasets compared to both classical and neural baselines, demonstrating the effectiveness of feature‑wise contrastive consistency for unsupervised feature selection.
By Utsab Ghosh, Roshni Chakraborty
CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.
By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie
arXiv:2105. 07610v5 Announce Type: replace-cross Abstract: Building trustworthy machine learning algorithms for biological applications requires adapting to data heterogeneity from different sources, batches, distributions, or studies.
By Maya Ramchandran, Rajarshi Mukherjee, Giovanni Parmigiani
arXiv:2606. 04665v1 Announce Type: new Abstract: Deep unsupervised domain adaptation (Deep UDA) methods successfully leverage rich labeled data in a source domain to boost the performance on related but unlabeled data in a target domain.
By Kaichao You, Ximei Wang, Mingsheng Long, Michael I. Jordan
arXiv:2606. 05139v1 Announce Type: new Abstract: The rapid advancement of high-throughput sequencing has led to large, high-dimensional omics datasets.
By Luca Thale-Bombien, Jan Ewald, Ralf K\"onig, Aaron Klein
arXiv:2603. 15263v2 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) has revolutionized representation learning, with Joint-Embedding Architectures (JEAs) emerging as an effective approach for capturing semantic features.
By Konstantinos Almpanakis, Anna Kreshuk
arXiv:2609.14882v1 Announce Type: cross
Abstract: Nucleotide sequence analysis is central to problems spanning regulatory genomics, evolutionary biology, and phenotype prediction. Classical bioinform...
By Evgeny S. Saveliev, Krzysztof Kacprzyk, Charlotte Capitanchik, Neelanjan Mukherjee, Kate Matlin, Ryan Sheridan, Srinivas Ramachandran, Jernej Ule, David L. Bentley, Mihaela van der Schaar
arXiv:2606. 13007v1 Announce Type: cross Abstract: Clustering is fundamental to scRNA-seq analysis, serving as a cornerstone for identifying cell populations and resolving tissue heterogeneity.
By Ping Xu, Pengjiang Li, Tian Du, Zaitian Wang, Jiawei Gu, Ziyue Qiao, Pengfei Wang, Yuanchun Zhou
The paper introduces a selective inference framework tailored for deep clustering that uses a fixed pretrained encoder to map high‑dimensional data into a latent space before clustering. It addresses the complex selection bias arising from the nonlinear transformation and offers a computationally tractable method to perform valid statistical tests on cluster differences. Experiments on synthetic data show controlled Type I error and higher power compared to conservative baselines, while genomic case studies demonstrate the ability to uncover significant cluster differences while properly accounting for selection bias.
By Eina Mizui, Tomohiro Shiraishi, Shunichi Nishino, Ichiro Takeuchi