arXiv Machine Learning By Aditya Upadhyay

UNIQ: Conformal Calibration for Adaptive Conservatism in Offline Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2606. 07592v1 Announce Type: new Abstract: Offline reinforcement learning requires careful conservatism to mitigate distribution shift, yet most existing methods apply a fixed penalty uniformly across all states regardless of local data coverage.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 29

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional wisdom suggests it should, but recent results show that online RL with a randomly-initialized Q-function can result in highly performant and reliable policies without needing to pretrain the Q-function.

arXiv Machine Learning
Jul 30

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

arXiv:2607. 27203v1 Announce Type: new Abstract: Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too?

By Perry Dong, Ron Polonsky, Dorsa Sadigh, Chelsea Fin
arXiv Machine Learning
Sep 18

Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation

The paper introduces Reward Stimulation Implicit Q-Learning (RSIQL), a non-hierarchical approach to improve offline goal-conditioned reinforcement learning. RSIQL adds auxiliary reward signals at intermediate states that are predicted to aid progress toward the goal, thereby reducing the delay in training supervision. Experiments on D4RL goal-reaching benchmarks and OGBench demonstrate that RSIQL outperforms baseline goal-conditioned IQL and rivals hierarchical offline methods while maintaining a simple flat policy structure.

By Jing Zhang