arXiv Machine Learning By Lars van der Laan, Nathan Kallus

Fitted Occupancy-Ratio Evaluation without Bellman Completeness

Read the original on arXiv Machine Learning →

arXiv:2607. 05375v1 Announce Type: cross Abstract: Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.