arXiv Machine Learning

Efficient Mean Curvature Computation on High-Dimensional Data Manifolds

arXiv:2606. 06329v1 Announce Type: new Abstract: Estimating local mean curvature at each point of a high-dimensional dataset is a key ingredient of geometry-aware machine learning algorithms, such as the Mean Curvature Boundary Points (MCBP) method.

arXiv Machine Learning
Aug 31

Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification

Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification (CARSANN) is a geometry-driven framework that adapts the spatial support of each neighborhood based on local geometric complexity. It estimates intrinsic dimensionality with TwoNN, builds an intrinsic representation via PCA, and uses a shape-operator-based estimate of local mean curvature to shrink the radius in highly curved regions while keeping a broader support in flatter areas. Experiments on over 70 OpenML datasets show that CARSANN consistently outperforms standard k‑NN and rivals other adaptive nearest‑neighbor methods, achieving a mean balanced accuracy increase from 0.6506 to 0.7528 and statistically significant improvements on most datasets.

By Alexandre L. M. Levada
arXiv Machine Learning
Sep 15

Eigenvalue-Decomposition Cost Denoising as an Alternative to Predict-then-Optimize for Shortest-Path Problems

The paper proposes using eigenvalue decomposition (or PCA) to denoise noisy cost observations for shortest‑path problems, instead of the traditional predict‑then‑optimize approach. By projecting new cost vectors onto the top‑k eigenvectors of the training covariance matrix before running Dijkstra’s algorithm, the method can recover the true underlying costs. Experiments on a 5×5 grid benchmark show that choosing k equal to the true latent feature dimension (k=5) yields the best performance, outperforming the SPO+ method especially under high model misspecification.

By Henry Aldridge-Krawciw, Irene Aldridge
arXiv Machine Learning
Jun 19

Spectral DPPs via NEPv: A Scalable Continuous Relaxation of Determinantal MAP for Diversity-Aware Data Selection

arXiv:2606. 19411v1 Announce Type: new Abstract: Selecting a small, diverse, high-quality subset from a massive pool of candidates is a recurring primitive in modern machine learning -- data curation and coreset selection for training and fine-tuning large models, active-learning batch acquisition, prompt and exemplar selection for in-context learning, retrieval diversification, and experimental design.

By Richard Yi Da Xu