arXiv Machine Learning

Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems

arXiv:2607. 23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation.

arXiv AI
Sep 7

One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation

The paper presents a single pretrained diffusion traffic model that serves both as an ego motion planner and as a controllable generator of safety‑critical scenarios for autonomous driving. It introduces a Single‑Stream Dual‑Stream diffusion‑transformer decoder (SSDS) that fuses scene context via joint attention, improving closed‑loop performance on the nuPlan benchmark, and a training‑free guidance scheme called Decoupled Annealing Posterior Sampling with Energy (DAPSE) that injects arbitrary energy functions at inference time. Using the same model, the authors generate realistic long‑tail driving interactions—such as aggressive cut‑ins and lead‑vehicle braking—through inference‑time guidance, exposing failure modes in black‑box planners that standard benchmarks miss.

By Arka Pal, Rajesh Kumar, Hannes Eriksson, R\'emi Lacombe, Arvid Laveno Ling, Ankit Gupta, Maciej Wozniak
arXiv Machine Learning
Aug 28

Sequential Additivity in Distributionally Robust Ranking and Selection

The paper studies distributionally robust ranking and selection (DRR&S), where the goal is to identify the best alternative under input uncertainty by considering multiple plausible input distributions. It introduces the concept of sequential additivity, showing that efficient sampling should focus on a small, additive set of critical scenarios rather than a multiplicative number. The authors prove an algorithm‑independent lower bound on sampling, design an additive allocation (AA) procedure that meets this bound and achieves exponentially decreasing error probability, and extend the approach to a general additive allocation (GAA) framework that incorporates traditional R&S sampling rules.

By Zaile Li, Yuchen Wan, L. Jeff Hong
arXiv AI
4d ago

Distinguish or Homogenize: Last-Chance Policy Identification and Risk-Budgeted Recovery under Irreversible Resource Depletion

The paper introduces the Distinguish-or-Homogenize principle, where an agent can either spend resources to differentiate between latent fault models or alter the system state so that the remaining models share a common acceptable policy, eliminating further diagnosis. This leads to the Last-Chance Policy Identification (LCPI) framework, which evaluates correctness at the reached state rather than the initial one, and defines the Last Identifiable Margin (LIM) as the boundary between distinguishing and homogenizing. For deterministic diagnostic graphs, an Exact-LIM recursion is provided, while for noisy finite-horizon recovery the authors propose Risk-Budgeted Compatibility Planning (RBCP), which searches a compatibility-aware frontier under a hard worst-case failure constraint, demonstrating improved risk-feasible recovery in microservice and MiniGrid scenarios.

By Yibo Guo, Xiaodan Wang
arXiv AI
Aug 25

Learning from the Test: Self-Referential Differential Testing for Deep RL Agents

The paper introduces Delta, a two‑phase framework for testing deep reinforcement learning agents. In the first phase, the agent under test is evaluated for catastrophic failures while collecting decision‑making data. The second phase trains a challenger agent from this data using offline RL; comparing the challenger’s rewards to the original agent reveals optimality bugs, and Delta successfully uncovered thousands of such issues across multiple environments.

By Junda He, Jieke Shi, Zhou Yang, Mingfei Cheng, David Lo