arXiv:2606. 24991v1 Announce Type: cross Abstract: Model Predictive Control (MPC) is widely used in industrial and robotic systems for enforcing constraints and embedding domain knowledge through finite-horizon optimization-based planning.
By Shambhuraj Sawant, Akhil S Anand, Dirk Reinhardt, Sebastien Gros
The paper proposes a hybrid PID–Deep Reinforcement Learning (DRL) controller for industrial processes, addressing the limitations of traditional PID controllers in complex, non‑linear, multi‑input environments. Using the Industrial Benchmark (IB) to test DRL, the authors develop a multi‑objective reward function and employ a TD3 agent to discover optimal settings for the IB’s ‘Gain’ and ‘Shift’ parameters. These parameters are then fed into a tuned PID controller, yielding a system that combines the optimal performance and efficiency of DRL with the reliability of classical control.
By Zhengyang (Cissy), Gu, Joseph E. Hernandez, John Burtenshaw, Sean Scott, Thomas Cook, Chris Couch
Model Predictive Control (MPC) is widely used in industrial and robotic systems for enforcing constraints and embedding domain knowledge through finite-horizon optimization-based planning. However, despite these strengths, an MPC scheme typically does not yield optimal policies for sequential decision-making problems formulated as Markov Decision Processes (MDPs).
arXiv:2604. 21030v2 Announce Type: replace-cross Abstract: The integration of Model Predictive Control (MPC) and Reinforcement Learning (RL) has emerged as a promising paradigm for constrained decision-making and adaptive control.
By Mohsen Jalaeian Farimani, Roya Khalili Amirabadi, Davoud Nikkhouy, Malihe Abdolbaghi, Mahshad Rastegarmoghaddam, Shima Samadzadeh, Mahdi Ghane
arXiv:2608. 10777v1 Announce Type: new Abstract: Linear Quadratic Stochastic Optimal Control (LQ-SOC) establishes a fundamental framework for steering noisy dynamical systems and has recently gained renewed interest in the machine learning community.
By Bangyan Liao, Chenglei Yu, Yuchen Yang, Chuanrui Wang, Zhisheng Song, Peidong Liu, Tailin Wu
arXiv:2605. 04568v3 Announce Type: replace-cross Abstract: State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, population-based methods for planning, learned policy networks, or a combination of policy networks and planning.
By Jonathan Spieler, Sven Behnke
The paper introduces Solver-Gradient Guided Reinforcement Learning (SG‑RL), a method that augments standard RL with bounded gradients from a differentiable MPC solver to adapt cost‑function weights online. SG‑RL integrates solver‑gradient guidance into PPO through actor‑update scaling, policy loss, advantage estimation, and value‑function learning, achieving comparable or superior closed‑loop performance while requiring up to 70.6% fewer samples. Experiments on two autonomous racing platforms with intentional model mismatch demonstrate that SG‑RL outperforms both RL and gradient‑based policy learning baselines and generalizes zero‑shot to unseen environments.
By Baha Zarrouki, Arslan Thobani, Jasper Hoffmann, Mattia Piccinini, Rudolf Reiter, Felix Jahncke, S\'ebastien Gros, Davide Scaramuzza, Johannes Betz
The paper presents a zero‑shot model predictive control (MPC) approach for buildings that uses excitation‑based generalized transfer learning models. By pretraining on purposefully probed operational data from multiple source buildings, the authors demonstrate that these models achieve superior control performance in 32 simulated target buildings, outperforming both an online linear MPC and a PI controller. This method eliminates the need for target‑specific data, reducing setup cost and facilitating broader deployment of energy‑efficient MPC in the building sector.
By Fabian Raisch, Felix Koch, Zack Xuereb Conti, Christoph Goebel, Benjamin Tischler
EfficientTDMPC is a sample‑efficient, model‑based reinforcement learning method for continuous control that builds on the TD‑MPC family. It reduces estimation error by using an ensemble of dynamics models and averaging return estimates across models and rollout depths, and it can penalize uncertainty in the planner objective. Practical improvements such as fresher buffer data and reduced compute enable the algorithm to benefit from a higher update‑to‑data ratio, achieving state‑of‑the‑art sample efficiency on HumanoidBench‑Hard and DMC hard while matching state‑of‑the‑art on DMC easy.
By Thomas Evers, Cristian Meo, Wendelin Bohmer, Justin Dauwels, Yaniv Oren
arXiv:2606. 24039v1 Announce Type: cross Abstract: Robotics increasingly relies on GPUs for parallel simulation, large-scale learning, and neural-network inference.
By Gabriel Bravo-Palacios, Jianghan Zhang, Zachary Pestrikov, Brian Plancher, Thomas Lew
arXiv:2507.06625v4 Announce Type: replace-cross
Abstract: Model Predictive Control (MPC) enables reliable trajectory optimization under dynamics constraints, but often depends on accurate dynamics mo...
By Shizhe Cai, Zeya Yin, Jayadeep Jacob, Fabio Ramos
arXiv:2608.20858v1 Announce Type: cross
Abstract: Transportation networks, in particular multi-class transportation networks (i.e., networks with mixed vehicle types), are complex systems that are ch...
By Giray Onur, Azita Dabiri, Bart De Schutter