arXiv Machine Learning

Meta-Learning-Assisted Constraint Relaxation for Constrained Black-Box Optimization

arXiv Machine Learning
Jun 10

Discovering Interpretable Multi-Parameter Control Policies for Evolutionary Algorithms Using Deep Reinforcement Learning

arXiv:2606. 10129v1 Announce Type: new Abstract: While deep Reinforcement Learning (deep-RL) has been increasingly applied to parameter control in evolutionary algorithms, rigorous theoretical analysis of parameter control remains largely restricted to single-parameter settings, owing to the difficulty of deriving effective, interpretable multi-parameter policies amenable to formal study.

By Tai Nguyen, Phong Le, Carola Doerr, Nguyen Dang
arXiv Machine Learning
Sep 2

Accelerating Reinforcement Learning via MPC Solver-Gradient Guidance for Weights-varying MPC

The paper introduces Solver-Gradient Guided Reinforcement Learning (SG‑RL), a method that augments standard RL with bounded gradients from a differentiable MPC solver to adapt cost‑function weights online. SG‑RL integrates solver‑gradient guidance into PPO through actor‑update scaling, policy loss, advantage estimation, and value‑function learning, achieving comparable or superior closed‑loop performance while requiring up to 70.6% fewer samples. Experiments on two autonomous racing platforms with intentional model mismatch demonstrate that SG‑RL outperforms both RL and gradient‑based policy learning baselines and generalizes zero‑shot to unseen environments.

By Baha Zarrouki, Arslan Thobani, Jasper Hoffmann, Mattia Piccinini, Rudolf Reiter, Felix Jahncke, S\'ebastien Gros, Davide Scaramuzza, Johannes Betz
arXiv Machine Learning
Aug 19

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.

By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
arXiv Machine Learning
4d ago

Global Optimality for Constrained Exploration via Penalty Regularization

The paper introduces Policy Gradient Penalty (PGP), a single‑loop policy‑space method that enforces convex occupancy‑measure constraints via quadratic‑penalty regularization. PGP constructs pseudo‑rewards to estimate gradients of the penalized objective and uses the classical Policy Gradient Theorem, establishing smoothness and global last‑iterate convergence guarantees for an ε‑optimal constrained entropy value with ε‑bounded constraint violation. The authors validate PGP with ablations on a grid‑world benchmark and demonstrate scalability on two challenging continuous‑control tasks.

By Florian Wolf, Ilyas Fatkhullin, Niao He
arXiv AI
Jun 19

Oranits: Mission Assignment and Task Offloading in Open RAN-based ITS using Metaheuristic and Deep Reinforcement Learning

arXiv:2507. 19712v3 Announce Type: replace-cross Abstract: In this paper, we explore mission assignment and task offloading in an Open Radio Access Network (Open RAN)-based intelligent transportation system (ITS), where autonomous vehicles leverage mobile edge computing for efficient processing.

By Ngoc Hung Nguyen, Nguyen Van Thieu, Quang-Trung Luu, Anh Tuan Nguyen, Senura Wanasekara, Nguyen Cong Luong, Fatemeh Kavehmadavani, Van-Dinh Nguyen
arXiv AI
Aug 26

Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable Heterogeneous Agile Earth Observation Satellite Scheduling

The paper introduces a reinforcement‑learning‑guided evolutionary policy optimization framework for scheduling heterogeneous agile Earth observation satellites, addressing task selection, satellite assignment, and sequencing under diverse visibility windows, maneuvering constraints, energy use, and storage limits. It combines assignment‑based indirect encoding with decoder‑based cost evaluation to capture satellite‑dependent constraints while integrating task gain, energy savings, and load balance into a single utility metric. The resulting RLOSMEA algorithm uses reinforcement learning to select high‑level search operators, achieving higher weighted utility and more stable convergence than baseline metaheuristics across varied AEOS scenarios.

By He Wang, Junyu Wu, Hui Li, Yanjie Song, Witold Pedrycz, Liang Li
arXiv Machine Learning
Sep 23

Differentiable Policy Transport over Multi-Layer Network Feasibility Geometry

The paper introduces Network Feasibility Geometry Reinforcement Learning (NFG‑RL), a method that enforces multi‑layer network constraints—such as interference, power‑rate coupling, flow conservation, service chains, capacity, latency, and reliability—by transporting a proto‑policy through a differentiable feasibility map. By compiling heterogeneous constraints into typed residual blocks and using a variational transport operator, NFG‑RL ensures almost‑sure feasible execution and shapes exploration and gradients to respect active constraints. Experiments on two wireless‑edge surrogate environments show that NFG‑RL boosts feasible utility by 37.5–41.5 %, cuts raw‑action violations by 48.5–60.8 %, and reduces P99 delay by 57.0–75.5 % compared to leading baselines.

By Zuyuan Zhang, Zeyu Fang, Mahdi Imani, Nathaniel D. Bastian, Tian Lan
arXiv AI
Jun 11

When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control

arXiv:2605. 26418v2 Announce Type: replace-cross Abstract: A properly calibrated rule-based autoscaler can beat every one of six mainstream deep reinforcement learning (DRL) algorithms on cost across every workload we test - so when, if ever, does DRL actually help?

By Guilin Zhang, Chuanyi Sun, Kai Zhao, Shahryar Sarkani, John Fossaceca