arXiv AI By Masafumi Endo, Kohei Honda, Yuu Jinnai, Ryo Yonetani

Orienteering Problem with Uncertain Time-Varying Rewards: Framework and Benchmark for Everyday Service Robotics

Read the original on arXiv AI →

The paper introduces the Orienteering Problem with Uncertain Time‑Varying Rewards (OP‑UTVR), a new variant of the classic orienteering problem that allows agents to estimate and forecast reward dynamics from observations. Three planners with different planning horizons and online adaptivity are proposed, and theoretical performance bounds under reward stochasticity are derived. A mobile service robot benchmark is presented, and experiments show trade‑offs between planning horizon and adaptivity, highlighting the benefits of long‑horizon planning with online adaptation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 19

Flickering Multi-Armed Bandits

arXiv:2602. 17315v3 Announce Type: replace-cross Abstract: We introduce Flickering Multi-Armed Bandits (FMAB) to model sequential decision-making in environments with changing action availability, where accessibility of the next action is restricted to a subset dependent on the agent's current choice.

By Sourav Chakraborty, Amit Kiran Rege, Claire Monteleoni, Lijun Chen
Hugging Face Trending Papers
Jul 15

From Novice to Expert: Cost-Aware Bandits for Evolving Worker Performance in Crowdsensing

Mobile crowdsensing (MC) recruits mobile users to perform sensing tasks using their smartphones, enabling large-scale applications such as traffic monitoring and environmental sensing. A fundamental challenge is online worker recruitment under uncertainty, where the platform must learn workers' sensing performance while operating with a limited budget.