arXiv Machine Learning

Model Predictive Path Integral PID Control for Learning-Based Path Following

arXiv:2603. 29499v2 Announce Type: replace-cross Abstract: Classical proportional--integral--derivative (PID) control remains widely used in industrial control systems, while model predictive control (MPC) is actively studied to achieve higher performance for systems with nonlinear dynamics.

arXiv Machine Learning
Sep 22

Augmenting PID Control with Deep Reinforcement Learning: A Hybrid Approach to the Industrial Benchmark

The paper proposes a hybrid PID–Deep Reinforcement Learning (DRL) controller for industrial processes, addressing the limitations of traditional PID controllers in complex, non‑linear, multi‑input environments. Using the Industrial Benchmark (IB) to test DRL, the authors develop a multi‑objective reward function and employ a TD3 agent to discover optimal settings for the IB’s ‘Gain’ and ‘Shift’ parameters. These parameters are then fed into a tuned PID controller, yielding a system that combines the optimal performance and efficiency of DRL with the reliability of classical control.

By Zhengyang (Cissy), Gu, Joseph E. Hernandez, John Burtenshaw, Sean Scott, Thomas Cook, Chris Couch
Hugging Face Trending Papers
Jun 23

Solving Markov Decision Processes with Future Information via MPC

Model Predictive Control (MPC) is widely used in industrial and robotic systems for enforcing constraints and embedding domain knowledge through finite-horizon optimization-based planning. However, despite these strengths, an MPC scheme typically does not yield optimal policies for sequential decision-making problems formulated as Markov Decision Processes (MDPs).

arXiv AI
Aug 6

A Systematic Review and Taxonomy of Reinforcement Learning-Model Predictive Control Integration for Linear Systems

arXiv:2604. 21030v2 Announce Type: replace-cross Abstract: The integration of Model Predictive Control (MPC) and Reinforcement Learning (RL) has emerged as a promising paradigm for constrained decision-making and adaptive control.

By Mohsen Jalaeian Farimani, Roya Khalili Amirabadi, Davoud Nikkhouy, Malihe Abdolbaghi, Mahshad Rastegarmoghaddam, Shima Samadzadeh, Mahdi Ghane
arXiv Machine Learning
Sep 2

Accelerating Reinforcement Learning via MPC Solver-Gradient Guidance for Weights-varying MPC

The paper introduces Solver-Gradient Guided Reinforcement Learning (SG‑RL), a method that augments standard RL with bounded gradients from a differentiable MPC solver to adapt cost‑function weights online. SG‑RL integrates solver‑gradient guidance into PPO through actor‑update scaling, policy loss, advantage estimation, and value‑function learning, achieving comparable or superior closed‑loop performance while requiring up to 70.6% fewer samples. Experiments on two autonomous racing platforms with intentional model mismatch demonstrate that SG‑RL outperforms both RL and gradient‑based policy learning baselines and generalizes zero‑shot to unseen environments.

By Baha Zarrouki, Arslan Thobani, Jasper Hoffmann, Mattia Piccinini, Rudolf Reiter, Felix Jahncke, S\'ebastien Gros, Davide Scaramuzza, Johannes Betz
arXiv Machine Learning
Sep 14

Very Exciting: Zero-Shot Model Predictive Control of Buildings via Excitation-Based Generalized Transfer Learning Models

The paper presents a zero‑shot model predictive control (MPC) approach for buildings that uses excitation‑based generalized transfer learning models. By pretraining on purposefully probed operational data from multiple source buildings, the authors demonstrate that these models achieve superior control performance in 32 simulated target buildings, outperforming both an online linear MPC and a PI controller. This method eliminates the need for target‑specific data, reducing setup cost and facilitating broader deployment of energy‑efficient MPC in the building sector.

By Fabian Raisch, Felix Koch, Zack Xuereb Conti, Christoph Goebel, Benjamin Tischler
arXiv AI
Sep 17

EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control

EfficientTDMPC is a sample‑efficient, model‑based reinforcement learning method for continuous control that builds on the TD‑MPC family. It reduces estimation error by using an ensemble of dynamics models and averaging return estimates across models and rollout depths, and it can penalize uncertainty in the planner objective. Practical improvements such as fresher buffer data and reduced compute enable the algorithm to benefit from a higher update‑to‑data ratio, achieving state‑of‑the‑art sample efficiency on HumanoidBench‑Hard and DMC hard while matching state‑of‑the‑art on DMC easy.

By Thomas Evers, Cristian Meo, Wendelin Bohmer, Justin Dauwels, Yaniv Oren