arXiv Machine Learning

Supervised Quadratic Feature Analysis: Information Geometry Approach for Dimensionality Reduction

arXiv:2502. 00168v5 Announce Type: replace-cross Abstract: Supervised dimensionality reduction maps labeled data into a low-dimensional feature space while preserving class separation.

arXiv Machine Learning
Aug 27

Efficient Estimation of High Information Projections using Nearest Neighbours

The paper introduces a new dimensionality reduction technique that enhances nearest‑neighbour relationships to estimate high‑information projections. It constructs a matrix encoding local covariance via nearest‑neighbour pairs and shows that, under standard regularity conditions, this matrix consistently estimates the Density Information Matrix (DIM), a non‑parametric analogue of the Fisher Information Matrix. The authors also demonstrate the method’s practical usefulness for clustering and outlier detection.

By David P. Hofmeyr
arXiv Machine Learning
Sep 17

A General Kernel Framework for Non-CND Distance Measures Using |D|-Dimensional Sparse Landmark Embeddings

The paper introduces the Sparse Landmark Embedding (SLE) kernel, a new framework that removes the need for conditionally negative definite (CND) distance measures in kernel methods and Gaussian Processes. By embedding each input into a sparse feature vector using compactly supported bump functions centered at all training points, any standard positive semi-definite (PSD) kernel can be applied in this embedding space, guaranteeing PSD for arbitrary distance measures. The authors provide theoretical guarantees on PSD, sparsity, stability, and universal approximation, and show through experiments with geodesic and Wasserstein distances that the SLE kernel matches or surpasses domain-specific baselines in predictive accuracy and uncertainty quantification.

By Marcus M. Noack, Maher B. Alghalayini, Mark D. Risser
arXiv Machine Learning
Sep 4

Anisotropic View Distance Metric for High-Dimensional Data: Theory, Geometry, and Fast Computation

The paper introduces View distance, a novel metric that projects high‑dimensional data onto all pairwise two‑dimensional planes and sums the Euclidean distances across these projections. It satisfies metric axioms, couples features, suppresses redundancy, and captures anisotropic geometry. To make it scalable, the authors propose a plane‑selection strategy using iterative Maximum Weight Matching, reducing complexity from ω(n²) to ω(k) and demonstrating competitive performance on twelve datasets.

By Yiqun Zhang, Hou-biao Li
arXiv Machine Learning
Sep 23

Relative Wasserstein Angle and the Problem of the $W_2$-Nearest Gaussian Distribution

The paper introduces a geometric framework for measuring how far empirical datasets deviate from the Gaussian family using optimal transport theory. It defines two new quantities—the relative Wasserstein angle and the orthogonal projection distance—based on the cone structure of the relative translation invariant quadratic Wasserstein space, and shows that the usual moment‑matching Gaussian is not generally the $W_2$‑nearest Gaussian. Closed‑form expressions are derived for one‑dimensional and several location–scale families, while a numerical approximation is proposed for higher dimensions, with experiments demonstrating convergence, stability, and the angle’s robustness as a non‑Gaussianity indicator.

By Binshuai Wang, Peng Wei