arXiv Machine Learning

Beyond Black-Box Advice: Learning-Augmented Algorithms for MDPs with Q-Value Predictions

arXiv:2307. 10524v3 Announce Type: replace Abstract: We study the tradeoff between consistency and robustness in the context of a single-trajectory time-varying Markov Decision Process (MDP) with untrusted machine-learned advice.

arXiv AI
Sep 11

Safe Learning Under Irreversible Dynamics via Asking for Help

The paper presents an algorithm that lets a learning agent ask for help from a mentor and transfer knowledge between similar states, enabling safe and effective learning in Markov decision processes with irreversible dynamics and infinite state spaces. It proves that both regret and the number of mentor queries grow sublinearly over time, using a sequence of three reductions to achieve a general result. The work claims to be the first formal proof that an agent can achieve high reward while becoming self‑sufficient in an unknown, unbounded, high‑stakes environment without resets.

By Benjamin Plaut, Juan Li\'evano-Karim, Hanlin Zhu, Stuart Russell
Hugging Face Trending Papers
Jun 23

Solving Markov Decision Processes with Future Information via MPC

Model Predictive Control (MPC) is widely used in industrial and robotic systems for enforcing constraints and embedding domain knowledge through finite-horizon optimization-based planning. However, despite these strengths, an MPC scheme typically does not yield optimal policies for sequential decision-making problems formulated as Markov Decision Processes (MDPs).

arXiv AI
2d ago

Q-Learning for Reachability in MEC-Free MDPs

The paper introduces Quasar, a model‑free Q‑learning algorithm that guarantees asymptotic convergence for reachability objectives in Markov Decision Processes that are free of non‑terminal maximal end components (MECs). Unlike prior model‑based methods, Quasar does not estimate transition probabilities, reducing memory usage from O(|S|²|A|) to O(|S||A|). Experiments on the Quantitative Verification Benchmark Set show that Quasar converges to optimal policies with far fewer samples than existing state‑of‑the‑art model‑based approaches.

By Lu-Chin Chang, Suguman Bansal
arXiv Machine Learning
Jul 1

End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions

arXiv:2603. 23461v2 Announce Type: replace Abstract: We study reinforcement learning (RL) with linear function approximation in Markov Decision Processes (MDPs) satisfying \emph{linear Bellman completeness} -- a fundamental setting where the Bellman backup of any linear value function remains linear.

By Zakaria Mhammedi, Alexander Rakhlin, Nneka Okolo
arXiv Machine Learning
Jun 2

All Models are Wrong, Knowing Where is Useful: On Model Uncertainty in Reinforcement Learning

arXiv:2606. 01363v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) infers information about the environment from a learned dynamics model and bears the potential to address open problems such as data efficient and safe learning in robotics.

By Bernd Frauenknecht, Devdutt Subhasish, Artur Eisele, Friedrich Solowjow, Sebastian Trimpe