Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward. Yet, modern applications such as language model fine-tuning or scientific discovery demand diversity.
The paper proposes a Multi-Objective Reinforcement Learning framework for portfolio optimization that incorporates ratings from three ESG agencies, addressing the divergence in ESG rating methodologies. It couples this with a Preference Elicitation system using Gaussian Processes, allowing users to infer latent utility functions via pairwise comparisons of portfolios based on Sharpe ratios and ESG scores. Experiments with LLM-generated portfolio managers show that regional background influences preference weights, with European personas prioritizing ESG alignment and Texas personas favoring risk‑adjusted returns.
By Giovanni Dispoto, Marcello Restelli, Carmine Ventre
arXiv:2606. 03962v1 Announce Type: cross Abstract: Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward.
By Anthony GX-Chen, Ankit Anand, Gheorghe Comanici, Zaheer Abbas, Eser Ayg\"un, David Smalling, Shibl Mourad, Doina Precup, Andr\'e Barreto, Mark Rowland
arXiv:2509. 03456v2 Announce Type: replace-cross Abstract: Off-policy evaluation (OPE) and off-policy learning (OPL) are foundational for decision-making in offline contextual bandits.
By Imad Aouali, Otmane Sakhi
arXiv:2607. 28408v1 Announce Type: new Abstract: This thesis studies policy learning in interactive systems where an agent observes a context, selects an action from a very large set, and receives partial feedback.
By Imad Aouali
arXiv:2603. 27044v3 Announce Type: replace-cross Abstract: Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality and substantial functional redundancy inherent to the policy parameter space.
By Andrea Fraschini, Davide Tenedini, Riccardo Zamboni, Mirco Mutti, Marcello Restelli
arXiv:2605. 07724v2 Announce Type: replace-cross Abstract: Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed reward signal, the model tends to collapse onto a narrow set of outputs that over-optimize that objective.
By Ali Falahati, Mohammad Mohammadi Amiri, Kate Larson, Lukasz Golab
arXiv:2606. 14929v1 Announce Type: cross Abstract: Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models.
By Yan Dai, Negin Golrezaei, Patrick Jaillet
arXiv:2011. 02565v2 Announce Type: replace-cross Abstract: Temporal abstraction allows reinforcement learning agents to represent knowledge and develop strategies over different temporal scales.
By Anand Kamat, Doina Precup
The paper introduces a method that integrates action abstraction into policy optimization for reinforcement learning and generative flow networks. By iteratively identifying frequently used action subsequences in high‑reward trajectories and treating them as single high‑level actions, the approach expands the action space and improves sample efficiency. Experiments on synthetic and real‑world tasks show that this technique discovers diverse high‑reward states more effectively, especially on challenging exploration problems, and yields interpretable abstract actions that reflect the underlying reward structure.
By Oussama Boussif, L\'ena N\'ehale Ezzine, Joseph D Viviano, Micha{\l} Koziarski, Moksh Jain, Esmeralda S. Whitammer, Emmanuel Bengio, Rim Assouel, Yoshua Bengio
arXiv:2607. 09298v1 Announce Type: cross Abstract: We study general-utility Markov decision processes (GUMDPs) with risk-aware objectives.
By Pedro P. Santos, F\'abio Vital, Alberto Sardinha, Francisco S. Melo
arXiv:2607. 08971v1 Announce Type: new Abstract: The stochastic linear bandit, where actions are represented as vectors and rewards are linear, is a central paradigm for sequential decision making.
By Gautam Dasarathy, Vineet Gattani, Lalit Jain