← Back to all news
arXiv Machine Learning September 17, 2026 By Yury Kolomeytsev

A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 29

Trust Region Masking for Long-Horizon LLM Reinforcement Learning

arXiv:2512. 23075v5 Announce Type: replace-cross Abstract: Policy gradient methods for Large Language Models optimize a policy $\pi_\theta$ via a surrogate objective computed from samples of a rollout policy $\pi_{\text{roll}}$.

By Yingru Li, Jiacai Liu, Jiawei Xu, Yuxuan Tong, Ziniu Li, Qian Liu, Baoxiang Wang
llmsreinforcement-learning
More like this →
arXiv Machine Learning
Jun 19

Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL

arXiv:2605. 05481v2 Announce Type: replace Abstract: We revisit a classic "chicken-and-egg" problem in reinforcement learning: to safely improve a policy, the value function must be accurate on the state-visitation distribution of the updated policy.

By Dillon Sandhu, Ronald Parr
reinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Aug 17

Offline Deep Q* Estimation with Diffusion Models

arXiv:2608. 14401v1 Announce Type: cross Abstract: In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations.

By Xiaohong Chen, Yuling Jiao, Lican Kang, Jerry Zhijian Yang, Chen Zhong
diffusionreinforcement-learning
More like this →
arXiv Machine Learning
Aug 21

Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic

arXiv:2608. 19587v1 Announce Type: new Abstract: While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored.

By Zhiqiang Tan
reinforcement-learning
More like this →
arXiv Statistics ML
Sep 2

Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits

arXiv:2511.08097v2 Announce Type: replace-cross Abstract: We consider a general infinite horizon Heterogeneous Restless multi-armed Bandit (RMAB). Heterogeneity is a fundamental problem for many real...

By Dheeraj Narasimha, Nicolas Gast
reinforcement-learning
More like this →
Hugging Face Trending Papers
Aug 20

Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic

While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored. Existing analyses often rely on double-loop architectures and invoke a linear entropy penalty.

reinforcement-learning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea