arXiv:2605. 08475v3 Announce Type: replace-cross Abstract: In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass.
By Mingsong Yan, Dongyang Li, Charles Kulick, Sui Tang
arXiv:2606. 26975v1 Announce Type: cross Abstract: Empirical Bayes (EB) estimators can match the first-order asymptotic risk of maximum likelihood (ML) while behaving very differently at second order: recent excess mean squared error (XMSE) analysis shows that kernel-based EB estimation may be worse than ML when the kernel is poorly aligned with the true parameter.
By Minghao Chen, Jiale Zheng
The paper proposes three information‑theoretic criteria for selecting the most relevant basis functions in sparse Gaussian process regression, tailored to different levels of prior knowledge. Experiments on six UCI regression datasets and three basis families (HSGP, VFF, VISH) show that the no‑data criterion is a robust default, often outperforming simple truncation, while the data‑aware criteria yield significant improvements for HSGP. The study demonstrates that careful basis‑function selection can lead to better performance without increasing computational cost.
By Marnix Van Soom, Ivan De Boi
arXiv:2606. 25169v2 Announce Type: replace-cross Abstract: Sampling from an unnormalized target by reversing an Ornstein-Uhlenbeck diffusion requires the score of each noise-perturbed marginal.
By Alois Duston, Tan Bui-Thanh
Empirical Bayes (EB) estimators can match the first-order asymptotic risk of maximum likelihood (ML) while behaving very differently at second order: recent excess mean squared error (XMSE) analysis shows that kernel-based EB estimation may be worse than ML when the kernel is poorly aligned with the true parameter. This paper turns that diagnostic into a design principle.
The paper introduces an amortized learning framework for selecting bandwidths in kernel density estimation by optimizing the logarithmic score across a distribution of tasks. It uses a truncated-and-renormalized bounded-support formulation and affine standardization to achieve stable learning and transferability across different intervals. Experiments on Gaussian samples, a multi-family benchmark, and randomized Gaussian mixtures demonstrate that the learned selector outperforms traditional methods such as Silverman’s rule, Sheather–Jones, and least‑squares cross‑validation, especially for small or heterogeneous samples.
By Junyi Liang, Hailiang Du
arXiv:2608.21729v1 Announce Type: new
Abstract: Simulation-Based Inference (SBI) serves as a vital framework for parameter inference in scientific fields where simulators involve intractable likeliho...
By Yichen Zang, Song Liu, Jiun-Yi Lin
arXiv:2606. 25169v1 Announce Type: cross Abstract: Sampling from an unnormalized target by reversing an Ornstein--Uhlenbeck diffusion requires the score of each noise-perturbed marginal.
By Alois Duston, Tan Bui Tanh
arXiv:2607. 08202v1 Announce Type: new Abstract: Estimating original-space conditional expectations is central to value-driven recommender systems, including dwell time, GMV, and LTV forecasting.
By Mingyu Zhao, Zhaohan Li, Zhenxiong Miao, Xu Zhang, Dewei Leng, Yanan Niu, Kun Gai
arXiv:2608. 08826v1 Announce Type: new Abstract: Adaptive procedures must work without nuisance information an oracle may use, such as a gradient scale or smoothness index, and robust procedures may have to answer queries whose coordinate and inspection time are chosen only after the data are seen.
By Ibne Farabi Shihab, Adria Binte Habib
The paper introduces an adaptive fitting procedure for mixtures of product distributions in Gaussian regression with a spike‑and‑slab prior, directly minimizing reverse Kullback‑Leibler divergence on inclusion indicators and active coefficients. This method jointly refines component parameters and weights as the mixture grows, avoiding extra divergence penalties on unused latent coefficients. Empirical results on 250 simulated datasets show that mixtures reduce errors in inclusion probabilities, grouped support probabilities, and coefficient covariance compared to multistart mean‑field approaches, and that direct joint refinement outperforms augmented or restricted refinement at fixed mixture size.
By Hanqing Li, Yaroslav Golub, Xuewen Lu
arXiv:2606. 02909v1 Announce Type: cross Abstract: Gradient observations can substantially improve Gaussian process (GP) surrogates, particularly in high-dimensional settings where function evaluations are expensive.
By Hyunseok Seung, Matthias Katzfuss