arXiv Machine Learning By Ami Tavory, Noam Touitou, Tal Sarig, Frank Cheng, Ido Guy

Bandits via Additive Quantized Representations

Read the original on arXiv Machine Learning →

The paper introduces Residual Quantization (RQ) as a representation layer for contextual bandits, converting continuous contexts into discrete centroid assignments across multiple levels. This approach allows additive bandit algorithms to capture nonlinear reward structures while maintaining strictly bounded memory and efficient online updates. Experiments on 13 datasets show that RQ variants outperform non-RQ counterparts on 11 datasets and match advanced baselines like XGBoost and neural methods using up to 1000 times less memory.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 29

From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning

The paper introduces Q-Target Pretrained Transformers (QTPT), a method that replaces supervised behavior cloning with a Bellman-style Q‑target objective for in‑context reinforcement learning. QTPT retains the context‑conditioned Transformer architecture but learns to estimate action values using rewards and transitions from the context, rather than merely imitating offline actions. The authors provide theoretical analysis in stochastic linear bandits and finite‑horizon MDPs, demonstrating improved robustness to weak or suboptimal data, and empirically show gains over supervised pretraining on controlled RL benchmarks and extensions to D4RL Kitchen and AntMaze.

By Yichen Lin, Xuyuan Xiong, Xue Wang, Xiangfu Meng, Mike Mingcheng Wei, Tao Yao
arXiv AI
Sep 2

Bandits in Prod: Hyperparameter Optimization at Inference Time

The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.

By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine