PERRY: Policy Evaluation with Confidence Intervals using Auxiliary Data
arXiv:2507. 20068v2 Announce Type: replace Abstract: Off-policy evaluation (OPE) methods estimate the value of a new reinforcement learning (RL) policy prior to deployment.
arXiv:2507. 20068v2 Announce Type: replace Abstract: Off-policy evaluation (OPE) methods estimate the value of a new reinforcement learning (RL) policy prior to deployment.
arXiv:2603. 09344v3 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables data-efficient and safe policy learning without online exploration, but its performance often degrades under distribution shift.
In value-based reinforcement learning, improving the accuracy of policy evaluation has been shown to improve downstream policy optimization performance. The widely adopted family of approximations rel...
arXiv:2608. 10634v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL), which learns environment dynamics to generate synthetic experience, is a promising approach to sample-efficient decision making.
The paper introduces robust successor features, a method that extends the successor representation to handle uncertainty in both reward functions and transition kernels within linear Markov Decision Processes. It provides a theoretical bound on Generalized Policy Improvement that quantifies performance loss due to mismatched dynamics, and demonstrates the approach on grid-based benchmarks against prior methods that consider only reward or transition differences.
Limiting‑Kernel Q(λ) (LKQL) is an off‑policy value estimator that blends n‑step truncation with a long‑horizon approximation based on the limiting kernel. It maintains the computational efficiency of n‑step methods while improving policy evaluation accuracy, especially for long‑horizon tasks. The authors prove faster convergence of LKQL’s operator under aperiodicity and near‑on‑policy conditions, and demonstrate empirical gains on MuJoCo continuous‑control benchmarks.
The paper introduces Q-Target Pretrained Transformers (QTPT), a method that replaces supervised behavior cloning with a Bellman-style Q‑target objective for in‑context reinforcement learning. QTPT retains the context‑conditioned Transformer architecture but learns to estimate action values using rewards and transitions from the context, rather than merely imitating offline actions. The authors provide theoretical analysis in stochastic linear bandits and finite‑horizon MDPs, demonstrating improved robustness to weak or suboptimal data, and empirically show gains over supervised pretraining on controlled RL benchmarks and extensions to D4RL Kitchen and AntMaze.
arXiv:2606. 27766v1 Announce Type: cross Abstract: Offline reinforcement learning enables policy learning from fixed datasets without additional environment interaction, making it appealing for safety-critical applications where online exploration is costly or unsafe.
arXiv:2605. 05481v2 Announce Type: replace Abstract: We revisit a classic "chicken-and-egg" problem in reinforcement learning: to safely improve a policy, the value function must be accurate on the state-visitation distribution of the updated policy.
arXiv:2601. 22211v2 Announce Type: replace Abstract: Reinforcement learning (RL) with combinatorial action spaces remains challenging because feasible action sets are exponentially large and governed by complex feasibility constraints, making direct policy parameterization impractical.
arXiv:2608. 03069v1 Announce Type: new Abstract: Deep Q-Networks (DQNs) learn value functions through bootstrapped temporal-difference updates, where future returns are approximated using a greedy maximization over next-state action values.
arXiv:2608. 02958v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks much like one making clean progress, because imitation supplies no notion of progress.