arXiv:2606. 05230v1 Announce Type: cross Abstract: Selecting a clustering algorithm and its hyperparameters without labels is a common difficulty in engineering machine learning pipelines that work with unsupervised analysis of sensor, image, or process data.
By Mahdi Shamsi, Soosan Beheshti
arXiv:2604. 18801v2 Announce Type: replace Abstract: Scientific particle simulations in cosmology, molecular dynamics, and fluid dynamics produce large-scale datasets whose storage, movement, and analysis increasingly rely on lossy compression.
By Congrong Ren, Sheng Di, Katrin Heitmann, Franck Cappello, Hanqi Guo
arXiv:2606. 14592v1 Announce Type: cross Abstract: Clustering is widely used for exploratory analysis and scientific discovery, driving insights from market segmentation to biological data analysis, but its outputs can be difficult to interpret, audit, and reproduce as modern datasets become increasingly large and complex.
By Claire M. He, Genevera I. Allen
arXiv:2502. 08397v3 Announce Type: replace-cross Abstract: Clustering is a fundamental technique in data analysis and machine learning, used to group similar data points together.
By Anna Livia Croella, Veronica Piccialli, Antonio M. Sudoso
arXiv:2606. 00302v1 Announce Type: cross Abstract: Despite being ubiquitous in science, clustering remains a technique whose results are not quantitatively scrutinized via a framework.
By Siamak K. Sorooshyari, Manuel A. Rivas, Robert Tibshirani
arXiv:2105. 07610v5 Announce Type: replace-cross Abstract: Building trustworthy machine learning algorithms for biological applications requires adapting to data heterogeneity from different sources, batches, distributions, or studies.
By Maya Ramchandran, Rajarshi Mukherjee, Giovanni Parmigiani
arXiv:2608.29045v1 Announce Type: new
Abstract: Biclustering, or co-clustering, aims to discover coherent submatrices by grouping rows and columns of a data matrix simultaneously. This local two-dime...
By Paritosh Tiwari, I Navin Kumar, James C. Bezdek, Punit Rathore
InfoTaxa presents an information‑calibrated, label‑free clustering approach for fine‑grained visual taxonomy, using frozen pretrained visual embeddings and DNA as an audit signal. On the BIOSCAN‑5M dataset, the method achieves 0.79 AMI at family and 0.67 at genus, outperforming prior image baselines and matching oracle‑K and graph‑based methods. The study shows that while clustering efficiency recovers most image‑available information at higher taxonomic ranks, species‑level performance remains limited by both clustering and representation, with DNA adding significant predictive value.
By David Ahmedt-Aristizabal, Mohammad Ali Armin, Lars Petersson
arXiv:2607. 19620v1 Announce Type: cross Abstract: In this paper, we present SCPP (Soft Clustering Python Package), an open-source Python framework for soft clustering.
By Kiyan Rezaee, Morteza Ziabakhsh, Artin Bahrampour, Seyed Mohammad Ghoreishi, Asal Khaje, Ali Sajedifar, Manny Chalak, Ava Zerafatangiz, Sadegh Eskandari
arXiv:2509. 25289v4 Announce Type: replace-cross Abstract: Identifying an effective clustering algorithm for a given dataset remains a fundamental unsupervised learning issue.
By Mohammadreza Bakhtyari, Bogdan Mazoure, Renato Cordeiro de Amorim, Guillaume Rabusseau, Vladimir Makarenkov
arXiv:2512.16481v2 Announce Type: replace-cross
Abstract: Survival analysis encompasses a broad range of methods for analyzing time-to-event data, with one key objective being the comparison of survi...
By Nora M. Villanueva, Marta Sestelo, Luis Meira-Machado
arXiv:2512. 16558v3 Announce Type: replace Abstract: Clustering is a cornerstone of modern data analysis.
By Dani\"el Bot, Leland McInnes, Jan Aerts