arXiv AI

Learnable Weighting of Intra-Attribute Distances for Categorical Data Clustering with Nominal and Ordinal Attributes

arXiv:2607. 05464v1 Announce Type: cross Abstract: The success of categorical data clustering generally much relies on the distance metric that measures the dissimilarity degree between two objects.

arXiv Machine Learning
Sep 24

Binary Classification from Coupled Pairwise Labels

The paper introduces SD-Pcomp learning, a binary classification framework that jointly utilizes Similarity/Dissimilarity (SD) labels and Pairwise Comparison (Pcomp) labels from instance pairs. It proposes an objective function that can be decomposed into either an SD estimator plus ordering information or a Pcomp estimator plus pair-type information, thereby integrating complementary relational cues. Experiments on eight datasets demonstrate that combining both label types improves classification accuracy and AUC compared to using either alone or a simple convex combination.

By Tomoya Tate, Kosuke Sugiyama, Masato Uchida
arXiv Statistics ML
Sep 3

Multidimensional scaling of two-mode three-way asymmetric dissimilarities: finding archetypal profiles and clustering

The paper extends the h‑plot multidimensional scaling technique to handle three‑way asymmetric dissimilarities, enabling the extraction of archetypal profiles and clustering in a unified Euclidean space. It provides an explicit eigenvector‑based solution that avoids local minima, is scale‑invariant, computationally efficient, and includes a straightforward goodness‑of‑fit assessment. The method is benchmarked against existing models and demonstrated on a financial dataset, with all data and code made publicly available for reproducibility.

By Mireia Mollar-Gumbau, Aleix Alcacer, Rafael Benitez, Vicente J. Bolos, Irene Epifanio
arXiv AI
Sep 4

Mixed Data Clustering Survey and Challenges

The paper "Mixed Data Clustering Survey and Challenges" discusses how the rise of big data has made clustering of heterogeneous datasets—containing both numerical and categorical variables—particularly difficult for traditional methods. It highlights the importance of hierarchical and explainable algorithms for producing interpretable results that aid decision‑making. The authors propose a new clustering approach based on pretopological spaces and benchmark it against classical numerical clustering algorithms and existing pretopological methods to evaluate its performance in the big data context.

By Maxence Choufa, Clement Cornet, Guillaume Guerard, Sonia Djebali, Loup-No\'e Levy
arXiv Machine Learning
Jun 10

$k$-Nearest Neighbors in Gromov--Wasserstein Space

arXiv:2606. 10295v1 Announce Type: cross Abstract: The Gromov--Wasserstein (GW) distance provides a framework for comparing metric measure spaces, regardless of their underlying structure or geometry.

By Kaitlyn Hohmeier, Nicolas Fraiman, Caroline Moosmueller
arXiv Machine Learning
Sep 4

Anisotropic View Distance Metric for High-Dimensional Data: Theory, Geometry, and Fast Computation

The paper introduces View distance, a novel metric that projects high‑dimensional data onto all pairwise two‑dimensional planes and sums the Euclidean distances across these projections. It satisfies metric axioms, couples features, suppresses redundancy, and captures anisotropic geometry. To make it scalable, the authors propose a plane‑selection strategy using iterative Maximum Weight Matching, reducing complexity from ω(n²) to ω(k) and demonstrating competitive performance on twelve datasets.

By Yiqun Zhang, Hou-biao Li