Finding the Time to Think: Learning Planning Budgets in Real-Time RL
arXiv:2606. 26463v1 Announce Type: new Abstract: Deliberating takes time.
arXiv:2506. 14411v2 Announce Type: replace-cross Abstract: In standard reinforcement learning (RL) settings, the interaction between the agent and the environment is typically modeled as a Markov decision process (MDP), which assumes that the agent observes the system state instantaneously, selects an action without delay, and executes it immediately.
arXiv:2606. 26463v1 Announce Type: new Abstract: Deliberating takes time.
Deliberating takes time. In real-time settings, that time is not free.
arXiv:2607. 05064v1 Announce Type: new Abstract: Reinforcement learning in real world environments often suffers from severe performance degradation due to delayed feedback.
arXiv:2603. 03480v2 Announce Type: replace Abstract: We study reinforcement learning with delayed state observation, where the agent observes the current state after some random number of time steps.
arXiv:2608. 11511v1 Announce Type: cross Abstract: In sequential decision making, an agent typically observes its environment and acts at every timestep.
arXiv:2607. 16421v1 Announce Type: new Abstract: It has long been recognized that humans have the ability to switch between fast, reactive decision-making and slower, deliberative planning.
arXiv:2607. 07252v1 Announce Type: new Abstract: Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics.
arXiv:2606. 18223v1 Announce Type: cross Abstract: With sophisticated cyber-attacks becoming increasingly prevalent, modern networks require intelligent autonomous cyber-defense agents trained via Reinforcement Learning (RL).
arXiv:2607. 19232v1 Announce Type: new Abstract: Hierarchical Reinforcement Learning (HRL) intends to separate strategic planning from primitive execution.
arXiv:2607. 26434v2 Announce Type: cross Abstract: Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematically widens the sim-to-real gap.
arXiv:2501. 12942v2 Announce Type: replace Abstract: Effective multi-user delay-constrained scheduling is crucial in various real-world applications, including embodied AI, instant messaging, live streaming, and data center management, where efficient resource allocation is required among users with diverse delay sensitivities.
arXiv:2603. 16842v2 Announce Type: replace Abstract: Stochastic resetting -- intermittently returning a process to a fixed reference state -- has emerged as an effective mechanism for optimizing first-passage properties.