SMC-ES: Automated synthesis of formally verified control policies
arXiv:2607. 15003v1 Announce Type: new Abstract: The deployment of autonomous cyber-physical systems in safety-critical environments requires closed-loop control strategies (i.
arXiv:2607. 23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation.
arXiv:2607. 15003v1 Announce Type: new Abstract: The deployment of autonomous cyber-physical systems in safety-critical environments requires closed-loop control strategies (i.
arXiv:2604. 23099v2 Announce Type: replace-cross Abstract: Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks.
arXiv:2608.22421v1 Announce Type: new Abstract: World models predict action-conditioned futures and serve as critical internal simulators for downstream planning and control. However, catastrophic pr...
arXiv:2608.21423v1 Announce Type: cross Abstract: Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonstrations to...
arXiv:2606. 31114v1 Announce Type: new Abstract: Unmanned Traffic Management (UTM) systems are cloud-based platforms designed to manage and coordinate multiple aerial vehicles remotely.
arXiv:2609.39995v2 Announce Type: new Abstract: Diffusion planners exhibit strong capabilities in generating multimodal trajectories. However, existing methods primarily rely on expert demonstrations...
arXiv:2608. 13719v1 Announce Type: new Abstract: Autonomous systems can fail in rare and heterogeneous ways, making real-world failure discovery difficult under limited testing budgets.
The paper presents a single pretrained diffusion traffic model that serves both as an ego motion planner and as a controllable generator of safety‑critical scenarios for autonomous driving. It introduces a Single‑Stream Dual‑Stream diffusion‑transformer decoder (SSDS) that fuses scene context via joint attention, improving closed‑loop performance on the nuPlan benchmark, and a training‑free guidance scheme called Decoupled Annealing Posterior Sampling with Energy (DAPSE) that injects arbitrary energy functions at inference time. Using the same model, the authors generate realistic long‑tail driving interactions—such as aggressive cut‑ins and lead‑vehicle braking—through inference‑time guidance, exposing failure modes in black‑box planners that standard benchmarks miss.
The paper studies distributionally robust ranking and selection (DRR&S), where the goal is to identify the best alternative under input uncertainty by considering multiple plausible input distributions. It introduces the concept of sequential additivity, showing that efficient sampling should focus on a small, additive set of critical scenarios rather than a multiplicative number. The authors prove an algorithm‑independent lower bound on sampling, design an additive allocation (AA) procedure that meets this bound and achieves exponentially decreasing error probability, and extend the approach to a general additive allocation (GAA) framework that incorporates traditional R&S sampling rules.
The paper introduces the Distinguish-or-Homogenize principle, where an agent can either spend resources to differentiate between latent fault models or alter the system state so that the remaining models share a common acceptable policy, eliminating further diagnosis. This leads to the Last-Chance Policy Identification (LCPI) framework, which evaluates correctness at the reached state rather than the initial one, and defines the Last Identifiable Margin (LIM) as the boundary between distinguishing and homogenizing. For deterministic diagnostic graphs, an Exact-LIM recursion is provided, while for noisy finite-horizon recovery the authors propose Risk-Budgeted Compatibility Planning (RBCP), which searches a compatibility-aware frontier under a hard worst-case failure constraint, demonstrating improved risk-feasible recovery in microservice and MiniGrid scenarios.
arXiv:2608. 04677v1 Announce Type: new Abstract: Algorithmic recourse seeks to help individuals reverse unfavorable automated decisions by recommending actionable changes that achieve a desired outcome.
The paper introduces Delta, a two‑phase framework for testing deep reinforcement learning agents. In the first phase, the agent under test is evaluated for catastrophic failures while collecting decision‑making data. The second phase trains a challenger agent from this data using offline RL; comparing the challenger’s rewards to the original agent reveals optimality bugs, and Delta successfully uncovered thousands of such issues across multiple environments.