arXiv Machine Learning By Xiaohong Chen, Yuling Jiao, Lican Kang, Jerry Zhijian Yang, Chen Zhong

Offline Deep Q* Estimation with Diffusion Models

Read the original on arXiv Machine Learning →

arXiv:2608. 14401v1 Announce Type: cross Abstract: In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.