arXiv Machine Learning

Rotation-Based Subspace Tracking for Robust Kernel PCA on Streaming Data

arXiv Machine Learning
Jul 15

Kernel PCA for Out-of-Distribution Detection: Non-Linear Kernel Selection and Approximation

arXiv:2505. 15284v2 Announce Type: replace Abstract: Out-of-Distribution (OoD) detection is vital for the reliability of deep neural networks, the key of which lies in effectively characterizing the disparities between OoD and In-Distribution (InD) data.

By Kun Fang, Qinghua Tao, Mingzhen He, Kexin Lv, Runze Yang, Haibo Hu, Xiaolin Huang, Jie Yang, Longbing Cao
arXiv Machine Learning
Sep 18

Online Supervised Dimension Reduction with Random Features: Diagnostics and Computational Trade-offs

The paper studies Online Kernel Supervised Principal Component Analysis (OKSPCA), which uses random features and an Adam-style orthonormal basis update to optimize a supervised spectral objective. It shows that accurate optimization of this objective does not guarantee accurate population subspace recovery or improved predictive performance, and it provides theoretical results on consistency, concentration, and perturbation of the estimator. Empirical experiments on six benchmarks reveal that replacing the tracker with the exact empirical target does not significantly change regression deficits, while classification-rank models capture most of the terminal objective energy but can exhibit substantial geometric deviation; sample-size studies further separate empirical accuracy from population recovery. The diagnostics also compare computational trade-offs, indicating that exact on-request computation can be faster in classification settings, whereas Adam saves time relative to full thin‑SVD in some dense regression requests, despite persistent geometric error.

By Zhenlin Yao, Wei Xiong
arXiv Machine Learning
Aug 20

Inference and Uncertainty Quantification for Streaming $r$-PCA

The paper tackles two key gaps in streaming PCA using Oja's algorithm: it establishes sharp operator‑norm convergence for general‑rank subspaces under sub‑Gaussian data, and it provides distributional inference for the resulting subspace estimator. The authors remove non‑vanishing remainder terms from existing analyses, achieving rates that match minimax bounds in both dense‑tail and sparse‑tail regimes. They further develop a linearization of Oja’s iterates, enabling high‑dimensional Gaussian approximations and an online multiplier bootstrap for practical inference.

By Haoshu Xu, Hongzhe Li
arXiv Machine Learning
Jun 16

Localized Kernel Projection Outlyingness: A Two-Stage Approach for Multi-Modal Outlier Detection

arXiv:2510. 24043v4 Announce Type: replace Abstract: This paper presents Two-Stage LKPLO, a novel multi-stage outlier detection framework that overcomes the coexisting limitations of conventional projection-based methods: their reliance on a fixed statistical metric and their assumption of a single data structure.

By Akira Tamamori
arXiv Machine Learning
Aug 26

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

The paper investigates diffusion models trained in a lazy high‑dimensional regime, extending benign overfitting theory to generative settings. By analyzing denoising score matching in a vector‑valued RKHS with an inner‑product kernel, the authors derive exact risk trajectories under gradient flow when the number of samples scales proportionally with dimensionality. These trajectories reveal three distinct phases—spectral generalization, noise‑dominated interpolation, and empirical Bayes memorization—whose interplay shapes the distribution of generated samples.

By Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz
arXiv Machine Learning
Jun 5

Anchor PCA

arXiv:2606. 06233v1 Announce Type: cross Abstract: Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques.

By Benedikt Seiter, Anya Fries, Julius von K\"ugelgen, Jonas Peters