arXiv Machine Learning By Youngmin Oh

Robust Linear Dueling Bandits with Post-serving Context under Unknown Delays and Adversarial Corruptions

Read the original on arXiv Machine Learning →

arXiv:2605. 01752v4 Announce Type: replace Abstract: We study linear dueling bandits in volatile environments characterized by the simultaneous presence of post-serving contexts, delayed feedback, and adversarial corruption.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.