The paper introduces ICOMT, a framework for interpretable clustering using optimal multi-way decision trees. It proposes a new discretization technique based on one-dimensional K‑means, formulates a binary linear optimization problem to ensure tree optimality, and demonstrates superior clustering accuracy and shallow tree structures on four public datasets.
By Hayato Suzuki, Shunnosuke Ikeda, Naoki Nishimura, Yuichi Takano
The paper "Mixed Data Clustering Survey and Challenges" discusses how the rise of big data has made clustering of heterogeneous datasets—containing both numerical and categorical variables—particularly difficult for traditional methods. It highlights the importance of hierarchical and explainable algorithms for producing interpretable results that aid decision‑making. The authors propose a new clustering approach based on pretopological spaces and benchmark it against classical numerical clustering algorithms and existing pretopological methods to evaluate its performance in the big data context.
By Maxence Choufa, Clement Cornet, Guillaume Guerard, Sonia Djebali, Loup-No\'e Levy
arXiv:2608. 05880v1 Announce Type: cross Abstract: Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data.
By Benjamin Connor, Anna Jurek-Loughrey, Lu Bai, Muhammad Fahim
arXiv:2605. 30225v2 Announce Type: replace Abstract: Clustering is an unsupervised technique for grouping data points by similarity.
By Pernille Matthews, Lena Krieger, Tommaso Amico, Artur Zimek, Thomas Seidl, Ira Assent
arXiv:2606. 14592v1 Announce Type: cross Abstract: Clustering is widely used for exploratory analysis and scientific discovery, driving insights from market segmentation to biological data analysis, but its outputs can be difficult to interpret, audit, and reproduce as modern datasets become increasingly large and complex.
By Claire M. He, Genevera I. Allen
arXiv:2511. 17823v2 Announce Type: replace Abstract: Clustering algorithms have long been the topic of research, representing the more popular side of unsupervised learning.
By Naitik Gada (Rochester Institute of Technology)
arXiv:2606. 28328v1 Announce Type: cross Abstract: In recent years, text clustering has become a critical technique for applications including intent discovery, topic mining, and recommendation systems.
By Daoming Wan, Yizheng Huang, Jimmy X. Huang
arXiv:2411. 01576v3 Announce Type: replace Abstract: The explainable clustering problem was first posed by Moshkovitz et al.
By Maximilian Fleissner, Maedeh Zarvandi, Debarghya Ghoshdastidar
The paper introduces absolute cluster indices that assess both compactness and separability of clusters, moving beyond relative measures commonly used in clustering validation. It defines a compactness function for each cluster and a set of neighboring points for cluster pairs to evaluate cluster quality and overall distribution margin. These indices are applied to determine the true number of clusters and are compared against widely-used validity indices on synthetic and real-world datasets.
By Adil M. Bagirov, Ramiz M. Aliguliyev, Nargiz Sultanova, Sona Taheri
arXiv:2606. 00302v1 Announce Type: cross Abstract: Despite being ubiquitous in science, clustering remains a technique whose results are not quantitatively scrutinized via a framework.
By Siamak K. Sorooshyari, Manuel A. Rivas, Robert Tibshirani
arXiv:2607. 19089v1 Announce Type: new Abstract: Breast cancer is one of the most widespread types of cancer, affecting approximately 8 million women worldwide.
By Davide Chicco, Nicoletta Benvenuto
arXiv:2607. 20799v1 Announce Type: new Abstract: Scalar metrics are often used to evaluate clusterings against known classes, but they can obscure a fundamental trade-off: clusterings should be informative about class labels while avoiding unnecessary fragmentation.
By Andreas Tiffeau-Mayer