← Back to all news
arXiv Machine Learning September 22, 2026 By Mehrdad Mohammadi, Qi Zheng, Ruoqing Zhu

Vector-Valued Distributional Reinforcement Learning Policy Evaluation: A Hilbert Space Embedding Approach

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • rag
  • reinforcement-learning
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Sep 22

Reinforcement learning with an expectile-based objective

arXiv:2602. 09300v2 Announce Type: replace Abstract: We consider the policy evaluation and control in a finite horizon reinforcement learning (RL) setting under an expectile-based objective.

By Shrey Rakeshkumar Patel, Sumedh Gupte, Soumen Pachal, Prashanth L. A., Sanjay P. Bhat
reinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Jun 18

Wasserstein Policy Learning for Distributional Outcomes

arXiv:2606. 19117v1 Announce Type: cross Abstract: Offline policy learning has received growing attention in causal inference.

By Yiyan Huang, Cheuk Hang Leung, Qi Wu, Zhiheng Zhang
More like this →
arXiv Machine Learning
Jun 11

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies

arXiv:2601. 08136v2 Announce Type: replace Abstract: Diffusion and flow policies are gaining prominence in online reinforcement learning (RL) due to their expressive power, yet training them efficiently remains a critical challenge.

By Zeyang Li, Sunbochen Tang, Navid Azizan
diffusionreinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Aug 17

Offline Deep Q* Estimation with Diffusion Models

arXiv:2608. 14401v1 Announce Type: cross Abstract: In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations.

By Xiaohong Chen, Yuling Jiao, Lican Kang, Jerry Zhijian Yang, Chen Zhong
diffusionreinforcement-learning
More like this →
Hugging Face Trending Papers
Jul 6

Fitted Occupancy-Ratio Evaluation without Bellman Completeness

Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class.

reinforcement-learning
More like this →
arXiv Machine Learning
Jun 9

Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning

arXiv:2605. 26078v3 Announce Type: replace Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions.

By Zhaoyu Zhu, Rui Gao, Shuang Li
diffusionreinforcement-learningsafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea