Proximal Residual Value Functions for Consistent Planning and Real-Time Execution
Read the original on arXiv Machine Learning →The paper introduces proximal residual value functions for two‑timescale decision systems, where a planning layer supplies a continuation‑value function to a real‑time optimizer that allocates resources, with inventory placement as a motivating example. The authors propose an end‑to‑end reinforcement learning method that learns a convex residual added to a strictly convex potential, enabling well‑posed optimization and end‑to‑end differentiation while maintaining an explicit convex objective for real‑time execution. They also provide necessary and sufficient conditions for smooth value functions to produce decisions consistent across planning and execution timescales, and demonstrate a 5.0% reduction in routing and transfer cost in an offline simulation using data from a large e‑commerce retailer.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.