arXiv Machine Learning

Balancing Optimality and Diversity: Human-Centered Decision Making through Generative Curation

arXiv:2409. 11535v3 Announce Type: replace Abstract: Many decision-support systems recommend actions by optimizing measurable objectives, even when a human decision-maker retains final authority and considers additional criteria that are difficult to specify in advance.

arXiv Machine Learning
Sep 3

Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization

The paper proposes a Multi-Objective Reinforcement Learning framework for portfolio optimization that incorporates ratings from three ESG agencies, addressing the divergence in ESG rating methodologies. It couples this with a Preference Elicitation system using Gaussian Processes, allowing users to infer latent utility functions via pairwise comparisons of portfolios based on Sharpe ratios and ESG scores. Experiments with LLM-generated portfolio managers show that regional background influences preference weights, with European personas prioritizing ESG alignment and Texas personas favoring risk‑adjusted returns.

By Giovanni Dispoto, Marcello Restelli, Carmine Ventre
arXiv AI
Jul 7

Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching

arXiv:2603. 27044v3 Announce Type: replace-cross Abstract: Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality and substantial functional redundancy inherent to the policy parameter space.

By Andrea Fraschini, Davide Tenedini, Riccardo Zamboni, Mirco Mutti, Marcello Restelli
arXiv AI
Jun 4

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences

arXiv:2605. 07724v2 Announce Type: replace-cross Abstract: Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed reward signal, the model tends to collapse onto a narrow set of outputs that over-optimize that objective.

By Ali Falahati, Mohammad Mohammadi Amiri, Kate Larson, Lukasz Golab
arXiv AI
Sep 3

Action abstractions for amortized sampling

The paper introduces a method that integrates action abstraction into policy optimization for reinforcement learning and generative flow networks. By iteratively identifying frequently used action subsequences in high‑reward trajectories and treating them as single high‑level actions, the approach expands the action space and improves sample efficiency. Experiments on synthetic and real‑world tasks show that this technique discovers diverse high‑reward states more effectively, especially on challenging exploration problems, and yields interpretable abstract actions that reflect the underlying reward structure.

By Oussama Boussif, L\'ena N\'ehale Ezzine, Joseph D Viviano, Micha{\l} Koziarski, Moksh Jain, Esmeralda S. Whitammer, Emmanuel Bengio, Rim Assouel, Yoshua Bengio