SCAPE is a scenario‑conditioned simulation‑augmented policy evaluation framework that predicts real‑world policy performance for specific scenarios using limited paired simulation‑and‑real samples and extensive simulation rollouts. It corrects sim‑to‑real bias in simulation labels before training the prediction model and calibrates prediction uncertainty via conformal prediction. Experiments on autonomous driving and quadruped velocity tracking show SCAPE reduces scenario‑level prediction error, improves testing sample efficiency, narrows calibrated prediction intervals, and generalizes better to out‑of‑distribution scenarios, enabling fine‑grained deployment strategies.
arXiv:2604. 09860v4 Announce Type: replace-cross Abstract: The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck due to rapid performance saturation and a lack of true generalization testing.
By Jenai Xuning Yang, Rishit Dagli, Alex Zook, Hugo Hadfield, Ankit Goyal, Stan Birchfield, Fabio Ramos, Jonathan Tremblay
arXiv:2606. 29898v1 Announce Type: cross Abstract: Real-world evaluation is the gold standard for robot policies because it tests them against the physical conditions and deployment challenges they are ultimately designed to handle.
By Haoxu Huang, Tongsam Zheng, Yifan Chen, Jiacheng You, Yang Gao
arXiv:2606. 10366v1 Announce Type: cross Abstract: Simulation has become an essential tool for evaluating and improving vision-language-action (VLA) policies, offering scalable, reproducible, and controllable alternatives to costly real-world robot evaluation.
By Shuo Wang, Hanyuan Xu, Yingdong Hu, Fanqi Lin, Yang Gao
arXiv:2606. 31106v1 Announce Type: cross Abstract: Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics.
By Hyeonchang Jeon, Kyungbeom Kim, Eugene Vinitsky, Kyung-Joong Kim
arXiv:2608.22549v1 Announce Type: new
Abstract: Batched simulators for autonomous driving have recently enabled training reinforcement learning (RL) agents at scale, encompassing thousands of traffic...
By Cevahir Koprulu, David Paz, Feng Tao, Yuliang Guo, Xinyu Huang, Ufuk Topcu, Liu Ren