SCAPE: Scenario-Conditioned Simulation-Augmented Policy Evaluation
arXiv:2608. 19425v1 Announce Type: cross Abstract: Reliable performance evaluation is a central bottleneck for deploying robot-learning policies in real-world conditions.
SCAPE is a scenario‑conditioned simulation‑augmented policy evaluation framework that predicts real‑world policy performance for specific scenarios using limited paired simulation‑and‑real samples and extensive simulation rollouts. It corrects sim‑to‑real bias in simulation labels before training the prediction model and calibrates prediction uncertainty via conformal prediction. Experiments on autonomous driving and quadruped velocity tracking show SCAPE reduces scenario‑level prediction error, improves testing sample efficiency, narrows calibrated prediction intervals, and generalizes better to out‑of‑distribution scenarios, enabling fine‑grained deployment strategies.
arXiv:2608. 19425v1 Announce Type: cross Abstract: Reliable performance evaluation is a central bottleneck for deploying robot-learning policies in real-world conditions.
arXiv:2606. 29898v1 Announce Type: cross Abstract: Real-world evaluation is the gold standard for robot policies because it tests them against the physical conditions and deployment challenges they are ultimately designed to handle.
arXiv:2606. 31106v1 Announce Type: cross Abstract: Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics.
arXiv:2604. 09860v4 Announce Type: replace-cross Abstract: The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck due to rapid performance saturation and a lack of true generalization testing.
arXiv:2606. 10366v1 Announce Type: cross Abstract: Simulation has become an essential tool for evaluating and improving vision-language-action (VLA) policies, offering scalable, reproducible, and controllable alternatives to costly real-world robot evaluation.
arXiv:2608.22549v1 Announce Type: new Abstract: Batched simulators for autonomous driving have recently enabled training reinforcement learning (RL) agents at scale, encompassing thousands of traffic...
arXiv:2607. 02431v1 Announce Type: cross Abstract: Reinforcement learning (RL) can overcome the demonstration-coverage limitation of imitation learning (IL) by allowing robots to improve through trial-and-error interaction beyond the states observed in demonstrations.
arXiv:2607. 04434v1 Announce Type: cross Abstract: Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities.
arXiv:2606. 27475v1 Announce Type: cross Abstract: Robots trained on real world data tend to be imprecise, slow, and brittle to perturbations.
arXiv:2607. 09866v1 Announce Type: cross Abstract: Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reproduction and diagnosis.
arXiv:2608. 10386v1 Announce Type: new Abstract: Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias.
arXiv:2602. 13977v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) promises to unlock capabilities beyond imitation learning for Vision--Language--Action (VLA) models, but its requirement for massive real-world interaction prevents direct deployment on physical robots.