This paper presents Chronocooked, a reinforcement learning (RL) benchmark suite for studying implicit interval timing in RL agents. Inspired by Overcooked, the suite comprises cooking scenarios that require temporal decision making.
arXiv:2607. 20656v1 Announce Type: cross Abstract: Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences.
By Manoosh Samiei, Doina Precup, Paul Masset
arXiv:2608. 11511v1 Announce Type: cross Abstract: In sequential decision making, an agent typically observes its environment and acts at every timestep.
By Christopher Watson, Arjun Krishna, Dinesh Jayaraman, Rajeev Alur
arXiv:2605. 11484v2 Announce Type: replace Abstract: Task completion in digital and physical environments increasingly involves complex temporal interaction, where actions and observations unfold over different time scales rather than align with fixed observation--action steps.
By Jialian Li, Yuchen Cao, Junhong Liu, Weiran Guo, Xutao Wang, Jiaming Song, Jiahao Zhang, Jie Chen
arXiv:2606. 26463v1 Announce Type: new Abstract: Deliberating takes time.
By Aneesh Muppidi, Firas Darwish, Dylan Cope, Jo\~ao F. Henriques, Jakob Nicolaus Foerster
Deliberating takes time. In real-time settings, that time is not free.
arXiv:2606. 19752v1 Announce Type: cross Abstract: Long-horizon robot manipulation policies trained with reward shaping can still exploit dense rewards through inefficient interaction, while rare efficient behaviors may be forgotten during training.
By Yinsen Jia, Boyuan Chen
arXiv:2608. 13625v1 Announce Type: new Abstract: Signal temporal logic (STL) provides a formal language for specifying real-time properties of real-valued observations, along with a quantitative robustness score for monitoring satisfaction.
By Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai
arXiv:2605. 19294v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) policies increasingly rely on asynchronous inference to hide large-model latency behind ongoing robot motion.
By Yixiang Zhu, Yonghao Chen, Zijie Yang, Yusong Hu, Xinyu Chen
arXiv:2511. 17855v5 Announce Type: replace Abstract: Robots must learn from both what people do and what they say, but either modality alone is often incomplete: physical corrections are grounded but ambiguous in intent, while language expresses high-level goals but lacks physical grounding.
By Jordan Abi Nader, David Lee, Nathaniel Dennler, Andreea Bobu
arXiv:2608. 08255v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments.
By Yifu Huo, Shunjie Xing, Chenglong Wang, Peinan Feng, Qiaozhi He, Yan Ding, Anxiang Ma, Yuxin Gao, Tongran Liu, Tong Xiao, Jingbo Zhu
arXiv:2512. 00062v2 Announce Type: replace-cross Abstract: Robotic policy learning for complex real-world manipulation tasks has seen rapid recent progress, enabled in large part by the ability to collect demonstrations through human operation.
By Taewook Nam, Junmo Cho, Youngsoo Jang, Sung Ju Hwang