Missing Data Imputation under Manifold Hypothesis
arXiv:2607. 03641v1 Announce Type: cross Abstract: The manifold hypothesis posits that high-dimensional data are concentrated near a low-dimensional embedded manifold.
arXiv:2607. 17390v1 Announce Type: cross Abstract: Kernel regression with tensor trains and Hadamard overparameterization (KReTTaH) is introduced as a training-data-free, interpretable, and nonparametric framework for multi-way data imputation.
arXiv:2607. 03641v1 Announce Type: cross Abstract: The manifold hypothesis posits that high-dimensional data are concentrated near a low-dimensional embedded manifold.
The paper introduces Coupled Tensor‑Tensor Completion (CTTC), a new framework that incorporates side information in tensor form to enhance tensor completion tasks. CTTC leverages hidden connections among multimodal tensors and is grounded in distance metric learning and group theory. Experiments on the DTD and LINCS datasets show that CTTC outperforms existing methods such as HaLRTC, CTRC, Cell, and NTDDR in both runtime and root‑sum‑of‑squares error for drug effect prediction.
arXiv:2412. 07041v4 Announce Type: replace-cross Abstract: Recovering incomplete multidimensional tensor-structured data is a fundamental task in many real-world applications.
The paper introduces Coupled Tensor‑Tensor Completion (CTTC), a new framework that incorporates side information in tensor form to enhance tensor completion tasks. CTTC leverages hidden connections among multimodal tensors and is grounded in distance metric learning and group theory. Experiments on the DTD and LINCS datasets show that CTTC outperforms existing methods such as HaLRTC, CTRC, Cell, and NTDDR in both run‑time and root‑sum‑of‑errors accuracy for predicting drug effects.
arXiv:2606. 16484v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold great potential for medicine, as they inherit knowledge from LLM and allow multiple data modalities to be integrated, analysed and interpreted in natural language.
arXiv:2606. 04176v1 Announce Type: new Abstract: We study a distributional generalization of the matrix completion problem in which each entry of the target matrix is a probability distribution rather than a scalar.
The paper introduces a functional Tucker decomposition (FTD) that incorporates a mode-wise continuity constraint into tensor factorization, modeling continuous modes as functions in a reproducing kernel Hilbert space (RKHS) without requiring a predefined basis. It preserves the multilinear subspace structure of the Tucker model and provides a reconstruction error bound for continuous modes, quantifying approximation quality when a subspace estimated on one domain is reused on another. The authors demonstrate the practical value of this subspace transfer on cross-domain classification tasks in hyperspectral imaging and multivariate time-series analysis.
arXiv:2607. 23295v1 Announce Type: cross Abstract: In real-world machine learning applications, incomplete observations create a fundamental challenge.
Tensor Completion using Subspace Information (TCSI) is an algorithm that leverages side information by estimating a subspace and reformulating tensor completion as a matrix regression problem. Theoretical analysis shows that accurate subspace information reduces sample complexity to nearly linear in the uncoupled ambient dimensions and relaxes signal-to-noise ratio requirements compared to existing guarantees. Numerical simulations and an application to reconstructing global Total Electron Content (TEC) maps demonstrate lower reconstruction errors than competing methods.
The paper introduces the Sparse Landmark Embedding (SLE) kernel, a new framework that removes the need for conditionally negative definite (CND) distance measures in kernel methods and Gaussian Processes. By embedding each input into a sparse feature vector using compactly supported bump functions centered at all training points, any standard positive semi-definite (PSD) kernel can be applied in this embedding space, guaranteeing PSD for arbitrary distance measures. The authors provide theoretical guarantees on PSD, sparsity, stability, and universal approximation, and show through experiments with geodesic and Wasserstein distances that the SLE kernel matches or surpasses domain-specific baselines in predictive accuracy and uncertainty quantification.
RDDMPI introduces a residual denoising diffusion model for multivariate time series imputation. By decomposing the missing signal into a baseline reconstruction and a residual uncertainty component, the method conditions the diffusion process on both the completed signal and its latent representation, using a reliability-aware mechanism to balance baseline influence. Experiments on benchmark datasets show that this approach improves reconstruction accuracy and uncertainty quantification compared to prior diffusion-based methods.
The paper introduces an online framework for functional principal component analysis (FPCA) tailored to multidimensional functional data streams. It models functional principal components with tensor product splines, enforcing smoothness and orthonormality via a penalized approach on a Stiefel manifold. The authors present efficient Riemannian stochastic gradient descent and AdaGrad algorithms, along with a dynamic smoothing parameter tuning strategy based on rolling block validation, and provide asymptotic normality results and confidence intervals for the estimators.