Hugging Face Trending Papers

Conformal Recovery-Deadline Certificates for Runtime Assurance of Adapting Controllers

Runtime assurance (RTA) protects a safety-critical system by switching from an advanced controller to a verified safe controller when a monitored condition is violated. The standard latching rule, which trips on the first breach of the safe set and then coasts, is correct for a diverging controller but pathological for a capable online-adapting one.

arXiv AI
Jul 2

Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems

arXiv:2607. 00334v1 Announce Type: new Abstract: Autonomous agents, whether LLM-driven software agents or robotic physical agents, face a common class of failure modes when operating without continuous human oversight: safety violations from unverified actions, behavioral instability from unconstrained loops, and continuity loss from unhandled error states.

By Srini Ramaswamy, Wang Miaosheng
arXiv AI
Sep 18

A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies

The paper introduces SPAR, a closed‑loop simulation platform that couples real‑time AUV control software with a higher‑level orchestration layer for fault injection, prompting, and evaluation of large language models (LLMs) in diagnosing and recovering from anomalies. SPAR enables ensemble testing of LLMs, comparing a frontier model with three locally deployable LLMs on a mass‑shift fault scenario across 480 trials, revealing that model choice significantly affects diagnostic accuracy. The study demonstrates that while the frontier model consistently ranks the correct fault mechanism among its top hypotheses, local models succeed mainly when they follow the full diagnostic procedure, and overall diagnosis and operational decisions appear decoupled in this dataset.

By Khalid Halba, Kylie Cooper, James G. Bellingham
arXiv AI
4d ago

Distinguish or Homogenize: Last-Chance Policy Identification and Risk-Budgeted Recovery under Irreversible Resource Depletion

The paper introduces the Distinguish-or-Homogenize principle, where an agent can either spend resources to differentiate between latent fault models or alter the system state so that the remaining models share a common acceptable policy, eliminating further diagnosis. This leads to the Last-Chance Policy Identification (LCPI) framework, which evaluates correctness at the reached state rather than the initial one, and defines the Last Identifiable Margin (LIM) as the boundary between distinguishing and homogenizing. For deterministic diagnostic graphs, an Exact-LIM recursion is provided, while for noisy finite-horizon recovery the authors propose Risk-Budgeted Compatibility Planning (RBCP), which searches a compatibility-aware frontier under a hard worst-case failure constraint, demonstrating improved risk-feasible recovery in microservice and MiniGrid scenarios.

By Yibo Guo, Xiaodan Wang
Hugging Face Trending Papers
Sep 17

A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies

The paper introduces SPAR, a closed‑loop simulation platform that couples real‑time AUV control software with a higher‑level orchestration layer to evaluate large language models (LLMs) for fault diagnosis and recovery. It demonstrates that a frontier LLM outperforms locally deployable models in identifying a mass‑shift fault, and shows that successful diagnosis depends on following a complete diagnostic procedure rather than premature conclusions. The study provides an architecture and ensemble evaluation methodology for LLM‑assisted mission management on low‑power autonomous underwater vehicles.

arXiv AI
Sep 25

Certified Task-Conditioned Active Observability

The paper introduces the concept of task‑conditioned active observability, defining the minimal interaction cost needed for an autonomous agent to identify task‑relevant states while guaranteeing safe abstention. It formalizes this complexity, proving that task‑predictive equivalence yields a unique minimal sufficient quotient that preserves complexity and eliminates unnecessary distinctions. The authors present theoretical characterizations for deterministic and noisy regimes, and demonstrate a certified observer that reduces sensor usage and model steps while maintaining zero false acceptances in extensive high‑dimensional trials.

By Linzhe Zhang, Changming Xu