arXiv Machine Learning

AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery

arXiv Machine Learning
Sep 4

Latent Energy Action Planning with World Models

Latent Energy Action Planning (LEAP) is a new method that treats the entire action horizon as a differentiable variable and optimizes it using a frozen LeWorldModel (LeWM). LEAP couples terminal latent goal matching with a terminal‑window state energy, ensuring that the predicted terminal latent and decoder‑predicted terminal descriptor align with the goal. Using a frozen goal‑conditioned proposal, a quasi‑Newton solver, and post‑optimization projection, LEAP improves mean success from 77.5% to 94.8% across four control domains while keeping the LeWM representation frozen.

By Phu Pham, Aniket Bera
arXiv Machine Learning
Aug 31

Meta-Prompt Optimization for LLM-Based Sequential Decision Making

The paper introduces EXPO, an algorithm that automatically optimizes the meta-prompt—specifically the task description and meta-instruction—for large language model agents in sequential decision-making tasks such as Bayesian optimization and multi-armed bandits. Building on adversarial bandit techniques to handle non-stationary rewards, the authors extend EXPO to EXPO-ES, which also optimizes exemplars (historical interactions) within the meta-prompt. Experiments demonstrate that these methods significantly improve the performance of LLM-based agents in sequential decision-making scenarios.

By Mingze Kong, Zhiyong Wang, Yao Shu, Zhongxiang Dai
arXiv Machine Learning
Sep 4

TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents

TIGPO (Temporal Instance-Graph Policy Optimization) extends graph-based credit assignment for long-horizon LLM agents by maintaining a persistent transition graph per task across policy updates. It allocates rollout budgets to both new exploration and revisiting past tasks, pairing current rollouts with earlier ones to create cross‑temporal references that stabilize advantage estimation. Experiments on ALFWorld and WebShop show TIGPO consistently outperforms previous group‑based and graph‑based policy optimization methods.

By Jinwei Gan
arXiv AI
Jun 18

STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability

arXiv:2606. 19236v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from policy entropy collapse during training.

By Haipeng Luo, Qingfeng Sun, Songli Wu, Can Xu, Wenfeng Deng, Han Hu, Yansong Tang