Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization
Read the original on arXiv Machine Learning →The paper proposes a Multi-Objective Reinforcement Learning framework for portfolio optimization that incorporates ratings from three ESG agencies, addressing the divergence in ESG rating methodologies. It couples this with a Preference Elicitation system using Gaussian Processes, allowing users to infer latent utility functions via pairwise comparisons of portfolios based on Sharpe ratios and ESG scores. Experiments with LLM-generated portfolio managers show that regional background influences preference weights, with European personas prioritizing ESG alignment and Texas personas favoring risk‑adjusted returns.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.