arXiv:2607. 16210v1 Announce Type: new Abstract: Reinforcement learning (RL) is increasingly applied in complex, safety-critical domains, yet the lack of rigorous behavioral guarantees for neural network-based policies remains a major barrier to deployment.
By Luca Marzari, Ezio Bartocci, Enrico Marchesini
arXiv:2606. 06123v1 Announce Type: new Abstract: When learning to walk, infants seem to address a coarse version of the problem first - stay upright, reach the caregiver - and refine it only when further practice at that resolution stops paying off.
By Fernando E. Rosas
arXiv:2606. 17377v1 Announce Type: new Abstract: We study performance-driven environment abstraction for decision-making in large Markov decision processes.
By Yue Guan, Dipankar Maity, Panagiotis Tsiotras
arXiv:2602. 06746v2 Announce Type: replace Abstract: We study multi-task reinforcement learning (RL), a setting in which an agent learns a single, universal policy capable of generalising to arbitrary, possibly unseen tasks.
By Alessandro Abate, Giuseppe De Giacomo, Mathias Jackermeier, Jan Kret\'insk\'y, Maximilian Prokop, Christoph Weinhuber
arXiv:2608. 13625v1 Announce Type: new Abstract: Signal temporal logic (STL) provides a formal language for specifying real-time properties of real-valued observations, along with a quantitative robustness score for monitoring satisfaction.
By Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai
The paper introduces a diagnostic workflow for multi‑objective reinforcement learning (MORL) that reveals behavioral differences among policies on the Pareto front, which are not apparent from value vectors alone. It offers quantitative and visual tools to inspect these variations and demonstrates their effectiveness on both simple grid tasks and more complex continuous‑control benchmarks.
By Antonio Mone, Zuzanna Osika, Florian Felten, Pradeep K. Murukannaiah, Mark Fuge, Frans A. Oliehoek, Luciano Cavalcante Siebert