arXiv AI By Qinchuan Cheng

Reserve-Aware Contrast Certificates for Conservative Bandits with Uncertain Baselines

Read the original on arXiv AI →

The paper introduces Reserve-C4B, a method for conservative bandits that ensures improvement over an incumbent policy while respecting a performance budget, even when the incumbent’s reward is uncertain. By focusing on the baseline-relative contrast and using a shared confidence set, the approach derives an exact expression for the avoidable penalty and a tighter admissibility test at each history. The method incorporates a reserve ledger to separate statistical evidence from performance deficit and a prefix-refresh extension to recertify decisions without discarding prior credit, achieving high-probability conditional-mean performance guarantees for linear rewards.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 18

Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits

The paper introduces Odds‑Ratio Thompson Sampling (OR‑TS), a method for batched multi‑armed bandits that updates the joint posterior over log‑odds contrasts and refits the shared level in each batch, rather than carrying over absolute reward rates. It presents a Bayesian bandit agent with controls for decay of past evidence and aggressiveness of allocation, and evaluates OR‑TS against traditional absolute‑rate memory across 86 public A/B series and synthetic environments. Results show that when the shared level varies significantly, OR‑TS outperforms absolute‑rate memory, reducing regret and ensuring the best arm receives more traffic, while also handling cases where contrasts themselves shift.

By Sulgi Kim
arXiv Machine Learning
Jun 2

Bandit Simulation for Average Reward Inference

arXiv:2606. 00913v1 Announce Type: cross Abstract: Multi-arm bandit algorithms are increasingly used in online platforms, clinical trials, and social science experiments, but valid statistical inference on their performance remains an open challenge.

By Samya Praharaj, Chih-Yu Chang, Koulik Khamaru, Kelly W. Zhang