arXiv Machine Learning By Nicolas Gutowski, Fabien Chhel, Alexandre Letard, Sylvain Lamprier

Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection

Read the original on arXiv Machine Learning →

arXiv:2607. 26273v1 Announce Type: new Abstract: We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of $k$ arms and observes their $d$-dimensional reward vectors under semi-bandit feedback.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.