arXiv:2609. 36142v1 Announce Type: cross Abstract: In Bayesian inference problems with non-Gaussian observation noise, the posterior is only as accurate as the noise density, and gradient-based samplers need that density and its gradient evaluable pointwise, whether from an explicit expression or from code, and without an inner solve.
By Joshua Chen, Peter Jan van Leeuwen
The paper tackles two key gaps in streaming PCA using Oja's algorithm: it establishes sharp operator‑norm convergence for general‑rank subspaces under sub‑Gaussian data, and it provides distributional inference for the resulting subspace estimator. The authors remove non‑vanishing remainder terms from existing analyses, achieving rates that match minimax bounds in both dense‑tail and sparse‑tail regimes. They further develop a linearization of Oja’s iterates, enabling high‑dimensional Gaussian approximations and an online multiplier bootstrap for practical inference.
By Haoshu Xu, Hongzhe Li
The paper develops a rigorous framework for computing the Normalized Maximum Likelihood (NML) codelength for regular path‑differentiable Lipschitz (PDL) estimators, which include non‑smooth models such as Lasso and Sparse SVMs. By leveraging geometric measure theory and a novel Propose‑and‑Project Metropolis‑Hastings sampler, the authors provide a method to exactly evaluate the stochastic complexity for these non‑smooth estimators and demonstrate its scalability to high‑dimensional settings. The study shows that the exact NML criterion can match cross‑validation performance while being more data‑efficient, offering a theoretically grounded alternative for model selection in modern machine learning.
By Trenton Lau, Gary P. T. Choi
arXiv:2602. 19126v2 Announce Type: replace Abstract: We propose a robust Bayesian formulation of random feature (RF) regression that accounts explicitly for prior and likelihood misspecification via Huber-style contamination sets.
By Michele Caprio, Katerina Papagiannouli, Siu Lun Chau, Sayan Mukherjee
arXiv:2606. 07289v1 Announce Type: new Abstract: Model merging combines several independently fine-tuned experts into a single multi-task model without any training data, reducing the storage, serving, and decentralized-development costs of large foundation models.
By Yongxian Wei, Runxi Cheng, Xingxuan Zhang, Li Shen, Chun Yuan, Peng Cui, Dacheng Tao
The paper tackles the challenge of predicting multiple high‑dimensional physical fields that must satisfy linear equality constraints, a common scenario in physics‑informed machine learning. It critiques the conventional approach of deducing one field from others, showing its sensitivity to arbitrary choices and its impact on accuracy and uncertainty. To address this, the authors introduce a symmetric framework that first applies a row‑wise PCA to preserve constraints in a latent space, then trains a linearly‑constrained multi‑output Gaussian process using a specially parametrized kernel, and validate the method on population dynamics and CFD problems involving Reynolds stress tensors.
By Mahamat Hamdan Nassouradine, Cl\'ement Gauchy, Pierre-Emmanuel Angeli, S\'ebastien da Veiga