arXiv Machine Learning By Gordei Verbii

Optimal Recovery Meets Bayesian Learning: Where Worst-Case Bounds Pay Off

Read the original on arXiv Machine Learning →

The paper shows that Worst‑Case Optimal Recovery (OR) and Bayesian learning solve the same Gaussian‑quadratic‑Hilbert problems, linking the radius of information to a nugget‑optimized Gaussian process posterior variance. It evaluates three Bayesian systems, demonstrating that OR can outperform Bayesian methods in certain calibration and reproducibility metrics, yet split‑conformal and other approaches can beat OR in interval scoring, especially under covariate shift. The authors propose matching the guarantee tool to the data regime and auditing that regime first.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 10

Sub-Quadratic Bisimulation Metrics via Approximate Nearest Neighbors: Coverage-Augmented Guarantees and Computable Two-Sided Certificates

arXiv:2608. 06762v1 Announce Type: new Abstract: Bisimulation metrics quantify behavioral similarity in Markov decision processes, but their Wasserstein fixed-point operator updates every state pair and incurs quadratic pairwise work.

By Ibne Farabi Shihab, Joyanta Jyoti Mondal
arXiv Machine Learning
Aug 19

The concentration game: Bayesian updating, regret, and information

The paper introduces a two-player zero-sum repeated game between a learner and nature that simultaneously captures Bayesian updating and an exact decomposition of exponential-weights regret. The game’s terminal payoff reflects the maximum gain a comparator can achieve given a fixed relative entropy from the prior, while the one-step constraint limits nature’s move by an information budget. The resulting regret splits into three precise components—per-round information loss, an additive retempering drift, and the comparator’s information relative to the prior—providing a unified framework that explains concentration phenomena, large-deviation bounds, and various learning methods such as bandits, posterior sampling, aggregation, and boosting.

By Akshay Balsubramani