← Back to all news
arXiv Machine Learning August 26, 2026 By Hwanwoo Kim, Panos Toulis, Eric Laber

Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • reinforcement-learning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Trending Papers
Aug 25

Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation

In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been propo...

reinforcement-learning
More like this →
arXiv Machine Learning
Aug 14

Variance Reduction Based Experience Replay for Policy Optimization

arXiv:2602. 05379v2 Announce Type: replace-cross Abstract: Effective reinforcement learning (RL) for complex stochastic systems requires leveraging historical data to improve sample efficiency and accelerate policy optimization.

By Hua Zheng, Wei Xie, M. Ben Feng, Keilung Choy
reinforcement-learningbenchmarkssafety
More like this →
arXiv Machine Learning
Jun 19

Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL

arXiv:2605. 05481v2 Announce Type: replace Abstract: We revisit a classic "chicken-and-egg" problem in reinforcement learning: to safely improve a policy, the value function must be accurate on the state-visitation distribution of the updated policy.

By Dillon Sandhu, Ronald Parr
reinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
6d ago

Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation

arXiv:2608.24146v1 Announce Type: new Abstract: In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate...

By Claire Chen, Shuze Daniel Liu, Licheng Luo, Rohan Chandra, Nan Jiang, Shangtong Zhang
reinforcement-learning
More like this →
arXiv Machine Learning
Jul 30

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning

arXiv:2607. 26509v1 Announce Type: new Abstract: Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement.

By Gong Gao, Xiao Lai, Ziqi Xie, Guojie Chen, Xianhui Liu, Weidong Zhao
reinforcement-learningsafety
More like this →
arXiv Machine Learning
Jun 16

Temporal Difference Learning for Diffusion Models

arXiv:2606. 15048v1 Announce Type: new Abstract: Diffusion models are typically trained with objectives that focus on local denoising targets at individual time steps (or adjacent pairs), which do not enforce consistency between predictions along the denoising trajectory.

By Qizhen Ying, Yangchen Pan, Victor Adrian Prisacariu, Junfeng Wen
diffusionreinforcement-learningbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea