arXiv Machine Learning

Reinforcement Learning-Guided NSGA-II Enhanced with Gray Relational Coefficient for Multi-Objective Optimization: Application to NASDAQ Portfolio Optimization

arXiv:2607. 16194v1 Announce Type: new Abstract: In modern financial markets, decision-makers increasingly rely on quantitative methods to navigate complex trade-offs among multiple, often conflicting objectives.

arXiv Machine Learning
Sep 3

Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization

The paper proposes a Multi-Objective Reinforcement Learning framework for portfolio optimization that incorporates ratings from three ESG agencies, addressing the divergence in ESG rating methodologies. It couples this with a Preference Elicitation system using Gaussian Processes, allowing users to infer latent utility functions via pairwise comparisons of portfolios based on Sharpe ratios and ESG scores. Experiments with LLM-generated portfolio managers show that regional background influences preference weights, with European personas prioritizing ESG alignment and Texas personas favoring risk‑adjusted returns.

By Giovanni Dispoto, Marcello Restelli, Carmine Ventre
arXiv AI
Jun 10

A Unified Multi-Modal Framework for Intelligent Financial Systems: Integrating Reinforcement Learning, High-Frequency Trading, and Game-Theoretic Approaches with Cross-Modal Sentiment Analysis

arXiv:2606. 10412v1 Announce Type: new Abstract: The rapid evolution of financial technology demands sophisticated artificial intelligence systems capable of handling diverse challenges across multiple domains simultaneously.

By Fanrong Liu, Zhang Yuwei, Mingni Luo
arXiv Machine Learning
Sep 21

Decision-Focused Learning for Mean-Variance Portfolio Optimization via KKT-Based Reformulation

The paper introduces a decision‑focused learning framework for mean‑variance portfolio optimization that embeds the Karush‑Kuhn‑Tucker optimality conditions of the lower‑level optimization into a single‑level learning problem. This approach preserves budget and short‑sale constraints while remaining tractable for standard nonlinear solvers. Experiments on real‑world ETF data across two asset universes demonstrate superior performance on multiple investment metrics and highlight the benefits of the proposed regularization.

By Kensei Nosaka, Shunnosuke Ikeda, Yuichi Takano
arXiv Machine Learning
Aug 19

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.

By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee