arXiv AI

Mixed Data Clustering Survey and Challenges

The paper "Mixed Data Clustering Survey and Challenges" discusses how the rise of big data has made clustering of heterogeneous datasets—containing both numerical and categorical variables—particularly difficult for traditional methods. It highlights the importance of hierarchical and explainable algorithms for producing interpretable results that aid decision‑making. The authors propose a new clustering approach based on pretopological spaces and benchmark it against classical numerical clustering algorithms and existing pretopological methods to evaluate its performance in the big data context.

arXiv AI
Jun 30

Interpretable Clustering: A Survey

arXiv:2409. 00743v4 Announce Type: replace-cross Abstract: In recent years, much of the research on clustering algorithms has primarily focused on enhancing their accuracy and efficiency, frequently at the expense of interpretability.

By Lianyu Hu, Mudi Jiang, Junjie Dong, Xinying Liu, Zengyou He
arXiv Machine Learning
Aug 24

Interpretable clustering via optimal multi-way decision trees

The paper introduces ICOMT, a framework for interpretable clustering using optimal multi-way decision trees. It proposes a new discretization technique based on one-dimensional K‑means, formulates a binary linear optimization problem to ensure tree optimality, and demonstrates superior clustering accuracy and shallow tree structures on four public datasets.

By Hayato Suzuki, Shunnosuke Ikeda, Naoki Nishimura, Yuichi Takano
arXiv Machine Learning
Jun 15

Cluster LOCO: Feature Importance For Interpreting Clusters

arXiv:2606. 14592v1 Announce Type: cross Abstract: Clustering is widely used for exploratory analysis and scientific discovery, driving insights from market segmentation to biological data analysis, but its outputs can be difficult to interpret, audit, and reproduce as modern datasets become increasingly large and complex.

By Claire M. He, Genevera I. Allen