arXiv AI By Christopher Watson, Arjun Krishna, Dinesh Jayaraman, Rajeev Alur

Let it Cook: Learning to Wait in Sequential Decision Making

Read the original on arXiv AI →

arXiv:2608. 11511v1 Announce Type: cross Abstract: In sequential decision making, an agent typically observes its environment and acts at every timestep.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jul 14

Adaptive Reinforcement Learning for Unobservable Random Delays

arXiv:2506. 14411v2 Announce Type: replace-cross Abstract: In standard reinforcement learning (RL) settings, the interaction between the agent and the environment is typically modeled as a Markov decision process (MDP), which assumes that the agent observes the system state instantaneously, selects an action without delay, and executes it immediately.

By John Wikman, Alexandre Proutiere, David Broman