arXiv Machine Learning
Jul 15

Kernel PCA for Out-of-Distribution Detection: Non-Linear Kernel Selection and Approximation

arXiv:2505. 15284v2 Announce Type: replace Abstract: Out-of-Distribution (OoD) detection is vital for the reliability of deep neural networks, the key of which lies in effectively characterizing the disparities between OoD and In-Distribution (InD) data.

By Kun Fang, Qinghua Tao, Mingzhen He, Kexin Lv, Runze Yang, Haibo Hu, Xiaolin Huang, Jie Yang, Longbing Cao
arXiv Machine Learning
Sep 18

Online Supervised Dimension Reduction with Random Features: Diagnostics and Computational Trade-offs

The paper studies Online Kernel Supervised Principal Component Analysis (OKSPCA), which uses random features and an Adam-style orthonormal basis update to optimize a supervised spectral objective. It shows that accurate optimization of this objective does not guarantee accurate population subspace recovery or improved predictive performance, and it provides theoretical results on consistency, concentration, and perturbation of the estimator. Empirical experiments on six benchmarks reveal that replacing the tracker with the exact empirical target does not significantly change regression deficits, while classification-rank models capture most of the terminal objective energy but can exhibit substantial geometric deviation; sample-size studies further separate empirical accuracy from population recovery. The diagnostics also compare computational trade-offs, indicating that exact on-request computation can be faster in classification settings, whereas Adam saves time relative to full thin‑SVD in some dense regression requests, despite persistent geometric error.

By Zhenlin Yao, Wei Xiong
arXiv Machine Learning
Aug 20

Inference and Uncertainty Quantification for Streaming $r$-PCA

The paper tackles two key gaps in streaming PCA using Oja's algorithm: it establishes sharp operator‑norm convergence for general‑rank subspaces under sub‑Gaussian data, and it provides distributional inference for the resulting subspace estimator. The authors remove non‑vanishing remainder terms from existing analyses, achieving rates that match minimax bounds in both dense‑tail and sparse‑tail regimes. They further develop a linearization of Oja’s iterates, enabling high‑dimensional Gaussian approximations and an online multiplier bootstrap for practical inference.

By Haoshu Xu, Hongzhe Li