Hugging Face Trending Papers

Dynamic Minimax Regret Optimization for Robust LLM Post-Training

Read the original on Hugging Face Trending Papers →

The paper introduces DUCB-OGD, an algorithm that couples a Discounted Upper‑Confidence‑Bound sampler with Online Gradient Descent to address dynamic minimax regret in robust large‑language‑model post‑training. It operates under instantaneous mini‑batch‑only bandit feedback, tracking worst‑source performance without re‑evaluating historical data. Experiments on fine‑tuning, preference optimization, and reinforcement learning demonstrate that DUCB‑OGD improves worst‑group robustness with negligible computational overhead.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 2

Bandits in Prod: Hyperparameter Optimization at Inference Time

The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.

By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
arXiv AI
Jun 9

Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

arXiv:2606. 09802v1 Announce Type: cross Abstract: We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized preference vector, and in the presence of context distributions that are drifting over time.

By Udvas Das, Waris Radji, Debabrota Basu, Odalric-Ambrym Maillard