Retraction-Free Optimization over the Stiefel Manifold for the LoRA Fine-Tuning
arXiv:2607. 25299v1 Announce Type: cross Abstract: Optimization over the Stiefel manifold plays a significant role in various machine learning tasks.
The paper introduces a geometry-aware Bayesian fine‑tuning method that uses Stein variational gradient descent on the Stiefel manifold. By transporting low‑rank adapter matrices along this manifold, the approach preserves orthogonality constraints and yields multiple inference solutions, enabling uncertainty quantification. Experiments demonstrate improved model calibration and higher prediction accuracy compared to Euclidean‑space SVGD and related methods.
arXiv:2607. 25299v1 Announce Type: cross Abstract: Optimization over the Stiefel manifold plays a significant role in various machine learning tasks.
arXiv:2606. 29184v1 Announce Type: new Abstract: While Low-rank adaptation (LoRA) enables highly efficient fine-tuning by constraining task-specific updates to fixed low-rank subspaces, this rigid design limits representational flexibility and often results in overconfident predictions and miscalibrated uncertainty, especially in low-data regimes.
arXiv:2606. 00413v1 Announce Type: cross Abstract: Sufficient dimension reduction (SDR) makes high-dimensional regression tractable by projecting the covariates onto a low-dimensional subspace that preserves the conditional mean of the response.
arXiv:2606. 16454v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) enables efficient adaptation of large pre-trained models to downstream tasks by parameterizing weight updates with low-rank matrices.
arXiv:2602.02848v2 Announce Type: replace Abstract: Advances in large language models have driven strong performance across many tasks, but their memory and compute costs still hinder deployment. SVD...
The paper investigates nonlinear dimensionality reduction for Bayesian optimisation (BO) by transforming high‑dimensional black‑box optimisation problems into a sequence of low‑dimensional latent‑space BO (LSBO) tasks. It extends earlier linear embedding approaches by using variational autoencoders (VAEs), deep metric loss, and adaptive retraining to better capture nonlinear structure, and couples LSBO with sequential domain reduction (SDR‑LSBO) to progressively narrow search domains. Experiments on GPU‑accelerated BoTorch with Matérn‑5/2 Gaussian‑process surrogates show that VAE‑based LSBO outperforms adaptive linear embeddings, and the authors provide a theoretical analysis of latent‑space error versus representation gap under PAC‑Bayes conditions.
arXiv:2601. 09361v4 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is a key paradigm for improving large-scale reasoning models.
The paper introduces a novel technique called "persistence of memory" to enhance stochastic subspace methods for large‑scale optimisation. By using a weakly correlated guidance vector that is refreshed only at wide intervals, the method provides a structured direction for random subspace descent. The authors demonstrate that this guidance can be efficiently computed in sparse or minibatch settings and present the first theoretical analysis of classical SSD methods for sparse functions, showing alignment with low‑lying Hessian eigenvectors near the optimum.
arXiv:2606. 12120v1 Announce Type: new Abstract: Low-rank optimal transport (OT) mitigates the quadratic scaling of classical solvers, yet existing approaches rely heavily on first-order mirror-descent updates that require careful hyperparameter tuning and ignore the optimization landscape's curvature.
arXiv:2609.21039v1 Announce Type: new Abstract: A pervasive structural pattern in modern deep learning is the linear factorization block: a submodule of the form $W = BA$ in which two parameter matri...
arXiv:2603. 07965v2 Announce Type: replace-cross Abstract: Bayesian optimization (BO) for high-dimensional constrained problems remains a significant challenge due to the curse of dimensionality.
arXiv:2607. 19498v1 Announce Type: cross Abstract: Gaussian process (GP) modeling is widely used in computational science and engineering.