Reinforcement learning with an expectile-based objective
arXiv:2602. 09300v2 Announce Type: replace Abstract: We consider the policy evaluation and control in a finite horizon reinforcement learning (RL) setting under an expectile-based objective.
arXiv:2602. 09300v2 Announce Type: replace Abstract: We consider the policy evaluation and control in a finite horizon reinforcement learning (RL) setting under an expectile-based objective.
arXiv:2606. 19117v1 Announce Type: cross Abstract: Offline policy learning has received growing attention in causal inference.
arXiv:2601. 08136v2 Announce Type: replace Abstract: Diffusion and flow policies are gaining prominence in online reinforcement learning (RL) due to their expressive power, yet training them efficiently remains a critical challenge.
arXiv:2608. 14401v1 Announce Type: cross Abstract: In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations.
Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class.
arXiv:2605. 26078v3 Announce Type: replace Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions.
Value functions are an essential component in actor-critic based deep reinforcement learning (RL). Conventionally, these functions are trained as a regression task by minimising the mean squared error (MSE) relative to bootstrapped target values.
arXiv:2603. 27044v3 Announce Type: replace-cross Abstract: Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality and substantial functional redundancy inherent to the policy parameter space.
arXiv:2607. 05375v1 Announce Type: cross Abstract: Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation.
arXiv:2607. 01880v1 Announce Type: new Abstract: Value functions are an essential component in actor-critic based deep reinforcement learning (RL).
arXiv:2501.06926v5 Announce Type: replace Abstract: Double reinforcement learning (DRL) provides efficient off-policy inference for policy values in nonparametric Markov decision processes (MDPs), bu...
arXiv:2606. 20206v1 Announce Type: cross Abstract: In offline Reinforcement Learning, immediate rewards in logged batch data are often unobserved due to sparse or irregular record-keeping, or censored beyond certain reward values.