arXiv Machine Learning

Nonlinear Factor Decomposition via Kolmogorov-Arnold Networks: A Spectral Approach to Asset Return Analysis

arXiv:2603. 28257v2 Announce Type: replace-cross Abstract: KAN-PCA is an autoencoder that uses a KAN as encoder and a linear map as decoder.

arXiv Machine Learning
Sep 3

Robust Streaming PCA

The paper studies streaming principal component analysis under a robust setting where the covariance matrix can vary within a temporal uncertainty set, rather than being fixed. It establishes fundamental convergence limits for any algorithm that recovers principal components and analyzes the noisy power method and Oja's algorithm, showing that the noisy power method achieves rate‑optimal convergence in this setting. Numerical experiments on synthetic and real‑world data confirm the theoretical findings.

By Daniel Bienstock, Minchan Jeong, Apurv Shukla, Se-Young Yun
arXiv Machine Learning
Jul 27

Heavy-Tailed Principal Component Analysis

arXiv:2603. 11308v3 Announce Type: replace Abstract: Principal Component Analysis (PCA) is a cornerstone of dimensionality reduction, yet its classical formulation relies critically on second-order moments and is therefore fragile in the presence of heavy-tailed data and impulsive noise.

By Mario Sayde, Christopher Khater, Jihad Fahs, Ibrahim Abou-Faycal
arXiv Machine Learning
Jun 30

Large and Deep Factor Models

arXiv:2402. 06635v3 Announce Type: replace-cross Abstract: We show that a deep neural network (DNN) trained to construct a stochastic discount factor (SDF) admits an additive decomposition separating nonlinear characteristic discovery from the pricing rule that aggregates them.

By Bryan Kelly, Boris Kuznetsov, Semyon Malamud, Yuan Zhang
arXiv Machine Learning
Aug 11

A solvable high-dimensional model where nonlinear autoencoders learn structure invisible to PCA while test loss misaligns with generalization

arXiv:2602. 10680v2 Announce Type: replace-cross Abstract: Many real-world datasets contain hidden structure that cannot be detected by simple linear correlations between input features.

By Vicente Conde Mendes, Lorenzo Bardone, C\'edric Koller, Jorge Medina Moreira, Vittorio Erba, Emanuele Troiani, Lenka Zdeborov\'a
arXiv Machine Learning
Sep 10

Temporal Kolmogorov-Arnold Networks (T-KAN) for High-Frequency Limit Order Book Forecasting: Efficiency, Interpretability, and Alpha Decay

The paper introduces Temporal Kolmogorov‑Arnold Networks (T‑KAN) for forecasting high‑frequency limit order book data, replacing fixed linear weights in LSTMs with learnable B‑spline activation functions. This approach captures the shape of market signals, yielding a 19.1% relative improvement in F1‑score at a 100‑step horizon and a 132.48% return versus a -82.76% drawdown for DeepLOB under 1.0 bps transaction costs. T‑KAN also offers interpretability through visible dead‑zones in the splines and is optimized for low‑latency FPGA deployment via High‑Level Synthesis.

By Ahmad Makinde
arXiv Machine Learning
Jul 15

Graph Regularized PCA

arXiv:2601. 10199v2 Announce Type: replace Abstract: Multivariate data often exhibit complex dependencies that violate the assumption of isotropic residual noise.

By Antonio Briola, Marwin Schmidt, Fabio Caccioli, Carlos Ros Perez, James Singleton, Christian Michler, Tomaso Aste
arXiv Statistics ML
Sep 4

Online Learning of Functional Principal Component Analysis for Multidimensional Functional Data

The paper introduces an online framework for functional principal component analysis (FPCA) tailored to multidimensional functional data streams. It models functional principal components with tensor product splines, enforcing smoothness and orthonormality via a penalized approach on a Stiefel manifold. The authors present efficient Riemannian stochastic gradient descent and AdaGrad algorithms, along with a dynamic smoothing parameter tuning strategy based on rolling block validation, and provide asymptotic normality results and confidence intervals for the estimators.

By Muye Nanshan, Nan Zhang, Jiguo Cao