← Back to all news
Hugging Face Trending Papers August 25, 2026

Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation

Read the original on Hugging Face Trending Papers →

The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.

  • reinforcement-learning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
6d ago

Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation

arXiv:2608.24146v1 Announce Type: new Abstract: In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate...

By Claire Chen, Shuze Daniel Liu, Licheng Luo, Rohan Chandra, Nan Jiang, Shangtong Zhang
reinforcement-learning
More like this →
arXiv AI
Jun 18

Robust Regularized Policy Iteration under Transition Uncertainty

arXiv:2603. 09344v3 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables data-efficient and safe policy learning without online exploration, but its performance often degrades under distribution shift.

By Hongqiang Lin, Zhenghui Fu, Weihao Tang, Pengfei Wang, Yiding Sun, Qixian Huang, Dongxu Zhang
reinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Aug 14

Variance Reduction Based Experience Replay for Policy Optimization

arXiv:2602. 05379v2 Announce Type: replace-cross Abstract: Effective reinforcement learning (RL) for complex stochastic systems requires leveraging historical data to improve sample efficiency and accelerate policy optimization.

By Hua Zheng, Wei Xie, M. Ben Feng, Keilung Choy
reinforcement-learningbenchmarkssafety
More like this →
arXiv AI
Jun 16

Safe Exploration via Policy Priors

arXiv:2601. 19612v3 Announce Type: replace-cross Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.

By Manuel Wendl, Yarden As, Manish Prajapat, Anton Pollak, Stelian Coros, Andreas Krause
agentsreinforcement-learningbenchmarkssafety
More like this →
arXiv Machine Learning
Jul 30

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning

arXiv:2607. 26509v1 Announce Type: new Abstract: Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement.

By Gong Gao, Xiao Lai, Ziqi Xie, Guojie Chen, Xianhui Liu, Weidong Zhao
reinforcement-learningsafety
More like this →
arXiv Machine Learning
6d ago

Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion

arXiv:2505.01361v3 Announce Type: replace Abstract: Temporal difference (TD) learning is a foundational algorithm in reinforcement learning (RL). For nearly forty years, TD learning has served as a w...

By Hwanwoo Kim, Panos Toulis, Eric Laber
reinforcement-learning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea