arXiv Machine Learning

Structured Adaptive Tensor Prediction for Streaming Data

arXiv:2606. 10085v1 Announce Type: new Abstract: Matrix-valued time series arise in a wide range of applications, such as spatio-temporal data from medical imaging and geophysics.

arXiv Machine Learning
Aug 19

Online Generalized Sparse Regression: How Does Overparametrization Help?

The paper introduces an online generalized-sparsity-constrained regression framework that addresses key challenges in online sparse regression, such as dynamic regularization, memory usage, and real-time computation. It proposes an efficient online hard‑thresholding algorithm that performs closed‑form updates using only summary statistics, achieving global convergence at optimal statistical rates when the projection set is overparameterized. Numerical experiments show the method consistently outperforms existing alternatives in online cardinality‑constrained linear regression and low‑rank matrix sensing.

By Shuoguang Yang, Qiang Sun
arXiv Machine Learning
5d ago

Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection

The paper introduces a Physics-informed Predictive Prior Tensor Autoencoder (PPPTAE) for anomaly detection in multi-dimensional time series. It combines reconstruction-based and prediction-based autoencoders by embedding a Bayesian fusion approach and a predictive prior that respects tensor correlations. The method incorporates physical laws via tensor low‑rank decomposition to prevent over‑generalization and is validated on real‑world datasets.

By Jianan Liu, Chunguang Li
arXiv Statistics ML
6d ago

Adaptive Subspace Modeling With Functional Tucker Decomposition

The paper introduces a functional Tucker decomposition (FTD) that incorporates a mode-wise continuity constraint into tensor factorization, modeling continuous modes as functions in a reproducing kernel Hilbert space (RKHS) without requiring a predefined basis. It preserves the multilinear subspace structure of the Tucker model and provides a reconstruction error bound for continuous modes, quantifying approximation quality when a subspace estimated on one domain is reused on another. The authors demonstrate the practical value of this subspace transfer on cross-domain classification tasks in hyperspectral imaging and multivariate time-series analysis.

By Noah Steidle, Joppe De Jonghe, Mariya Ishteva
arXiv Machine Learning
Jun 25

Tensorion: A Tensor-Aware Generalization of the Muon Optimizer

arXiv:2606. 25975v1 Announce Type: new Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models.

By Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Sergei Kudriashov, Maxim Rakhuba
arXiv Machine Learning
Jun 4

Low-rank Distributional Matrix Completion

arXiv:2606. 04176v1 Announce Type: new Abstract: We study a distributional generalization of the matrix completion problem in which each entry of the target matrix is a probability distribution rather than a scalar.

By Jiayi Wang, Raymond K. W. Wong
arXiv Machine Learning
5d ago

Online Learning via Learned Latent Bayesian Tracking

The paper introduces AURA, a meta‑learning framework that learns a low‑dimensional latent state‑space model for the evolution of optimal model parameters under distribution shift. Online adaptation is performed via extended Kalman filtering in this latent space, followed by reconstruction of full model parameters through a learned lifting map, enabling efficient single‑step updates. Experiments on neural wireless receivers and non‑stationary image classification show that AURA improves adaptation speed, accuracy, and computational efficiency compared to existing online learning and Bayesian filtering baselines.

By Guy Gerson, Tomer Raviv, Nir Shlezinger, Tirza Routtenberg, Osvaldo Simeone
arXiv AI
Sep 21

Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods

The paper introduces Row-wise Matrix AdaGrad and Column-wise Matrix AdaGrad, two adaptive subgradient methods that extend AdaGrad to matrix-valued parameters by using row-wise and column-wise proximal functions. It presents a general Online Mirror Descent framework that derives these optimizers through online regret minimization, providing regret guarantees that can be tighter than entry-wise AdaGrad for structured gradients. Experiments on matrix factorization and deep neural-network training show that aligning adaptive scaling with matrix structure improves optimization stability, allows larger learning rates, and supports greater network depth.

By Wenpeng Zhang, Runsheng Yu, Peilin Zhao