arXiv:2606. 26397v1 Announce Type: cross Abstract: Real-world decision-making often requires balancing multiple conflicting objectives, a challenge that standard Reinforcement Learning (RL) frequently addresses by aggregating rewards into a single scalar signal.
By Aniruddha Joshi, Niklas Lauffer, Sanjit Seshia
arXiv:2606. 18111v1 Announce Type: cross Abstract: Fairness is an important aspect of decision-making in multi-objective reinforcement learning (MORL), where policies must ensure both optimality and equity across multiple, potentially conflicting objectives.
By Umer Siddique, Peilang Li, Yongcan Cao
arXiv:2608. 02826v1 Announce Type: cross Abstract: Reinforcement learning is a subfield of machine learning that studies how an agent interacts with an environment in order to extract as large a reward as possible.
By Joao F. Doriguello
arXiv:2607. 28916v1 Announce Type: cross Abstract: Multistep credit assignment is critical for sample-efficient reinforcement learning, yet managing off-policy bias in Q-learning remains a fundamental challenge.
By Brett Daley
The paper introduces Reward Stimulation Implicit Q-Learning (RSIQL), a non-hierarchical approach to improve offline goal-conditioned reinforcement learning. RSIQL adds auxiliary reward signals at intermediate states that are predicted to aid progress toward the goal, thereby reducing the delay in training supervision. Experiments on D4RL goal-reaching benchmarks and OGBench demonstrate that RSIQL outperforms baseline goal-conditioned IQL and rivals hierarchical offline methods while maintaining a simple flat policy structure.
By Jing Zhang
The paper introduces a framework for combining large language models (LLMs) with reinforcement learning (RL) by treating the LLM as a planner and the RL agent as a controller. It formalizes this hybrid setup as a Goal-Augmented Markov Decision Process and proves that using the LLM’s per‑state progress score as a bounded potential function preserves the optimal policy set, even if the LLM scores are inaccurate. The authors validate their theoretical result with numerical experiments on a small MDP, testing four potential configurations, including an adversarial case with a potential scaled twenty times the base reward.
By Christophe D. Hounwanou, John Emeka Eze, Ya\'e U. Gaba