Hugging Face Trending Papers

Fitted Occupancy-Ratio Evaluation without Bellman Completeness

Read the original on Hugging Face Trending Papers →

Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.