STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation
Read the original on arXiv Machine Learning →STITCH-OPE is a model‑based generative framework that uses denoising diffusion to perform off‑policy evaluation (OPE) in high‑dimensional, long‑horizon settings. It generates synthetic trajectories for a target policy by guiding a diffusion model trained on behavior data, subtracting the behavior policy’s score to avoid over‑regularization and stitching partial trajectories to extend horizon length. The authors provide theoretical variance‑reduction guarantees and demonstrate improved mean squared error, correlation, and regret on D4RL and OpenAI Gym benchmarks.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.