← Back to all news
arXiv Machine Learning August 26, 2026 By Claire Chen, Shuze Daniel Liu, Licheng Luo, Rohan Chandra, Nan Jiang, Shangtong Zhang

Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • reinforcement-learning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Trending Papers
Aug 25

Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation

In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been propo...

reinforcement-learning
More like this →
arXiv AI
Jun 18

Robust Regularized Policy Iteration under Transition Uncertainty

arXiv:2603. 09344v3 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables data-efficient and safe policy learning without online exploration, but its performance often degrades under distribution shift.

By Hongqiang Lin, Zhenghui Fu, Weihao Tang, Pengfei Wang, Yiding Sun, Qixian Huang, Dongxu Zhang
reinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Aug 14

Variance Reduction Based Experience Replay for Policy Optimization

arXiv:2602. 05379v2 Announce Type: replace-cross Abstract: Effective reinforcement learning (RL) for complex stochastic systems requires leveraging historical data to improve sample efficiency and accelerate policy optimization.

By Hua Zheng, Wei Xie, M. Ben Feng, Keilung Choy
reinforcement-learningbenchmarkssafety
More like this →
arXiv Machine Learning
Jun 30

Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning

arXiv:2605. 29032v2 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss.

By Christoph Dann, Yishay Mansour, Mehryar Mohri
agentsreinforcement-learningsafety
More like this →
arXiv AI
Jun 16

Safe Exploration via Policy Priors

arXiv:2601. 19612v3 Announce Type: replace-cross Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.

By Manuel Wendl, Yarden As, Manish Prajapat, Anton Pollak, Stelian Coros, Andreas Krause
agentsreinforcement-learningbenchmarkssafety
More like this →
arXiv Machine Learning
Jul 30

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning

arXiv:2607. 26509v1 Announce Type: new Abstract: Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement.

By Gong Gao, Xiao Lai, Ziqi Xie, Guojie Chen, Xianhui Liu, Weidong Zhao
reinforcement-learningsafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea