Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems
arXiv:2607. 20005v1 Announce Type: new Abstract: In modern IT operations (IT-Ops), the cost of an incorrect repair often exceeds the cost of no action at all.
The paper introduces the Distinguish-or-Homogenize principle, where an agent can either spend resources to differentiate between latent fault models or alter the system state so that the remaining models share a common acceptable policy, eliminating further diagnosis. This leads to the Last-Chance Policy Identification (LCPI) framework, which evaluates correctness at the reached state rather than the initial one, and defines the Last Identifiable Margin (LIM) as the boundary between distinguishing and homogenizing. For deterministic diagnostic graphs, an Exact-LIM recursion is provided, while for noisy finite-horizon recovery the authors propose Risk-Budgeted Compatibility Planning (RBCP), which searches a compatibility-aware frontier under a hard worst-case failure constraint, demonstrating improved risk-feasible recovery in microservice and MiniGrid scenarios.
arXiv:2607. 20005v1 Announce Type: new Abstract: In modern IT operations (IT-Ops), the cost of an incorrect repair often exceeds the cost of no action at all.
The paper introduces a framework for diagnosing and recovering from hidden dynamics changes in deployed control policies, focusing on the problem of task readiness under dormant dynamics drift. It proposes an intervention-based Bayesian method called Evidence‑Gated Matched‑Pulse Transport that localizes faults and estimates actuator effectiveness, enabling agents to certify readiness for future tasks with limited, task‑agnostic interactions. The approach is evaluated on diverse benchmarks, measuring readiness coverage, selective risk, interaction cost, and return, and identifies regimes where transported evidence is decisive.
arXiv:2607. 19338v1 Announce Type: new Abstract: Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer.
arXiv:2607. 09770v1 Announce Type: new Abstract: Industrial agentic AI systems increasingly exhibit a gap between prototype capability and production deployment.
The paper introduces the concept of task‑conditioned active observability, defining the minimal interaction cost needed for an autonomous agent to identify task‑relevant states while guaranteeing safe abstention. It formalizes this complexity, proving that task‑predictive equivalence yields a unique minimal sufficient quotient that preserves complexity and eliminates unnecessary distinctions. The authors present theoretical characterizations for deterministic and noisy regimes, and demonstrate a certified observer that reduces sensor usage and model steps while maintaining zero false acceptances in extensive high‑dimensional trials.
arXiv:2609.20973v1 Announce Type: cross Abstract: Complex problem solving (CPS) with large language models (LLMs) is often framed as a matter of stronger reasoning or longer generation. Yet early-ste...
arXiv:2607. 23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation.
arXiv:2606. 28011v1 Announce Type: cross Abstract: We propose an agentic Large Language Model (LLM) framework for active Fault-Tolerant Control (FTC) that transforms fault detection outputs into constraint-aware recovery actions grounded in plant-specific knowledge.
Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap model first, then escalate hard cases to a stronger and more expensive model.
arXiv:2607. 18243v1 Announce Type: new Abstract: Agentic AI is crossing trust boundaries faster than current risk models can represent.
Self-correction is particularly useful when a failure constrains the next repair. Coding agents benefit from this property because compilers, tests, and execution traces turn many failures into typed recovery signals, but broad language-agent tasks often expose only a coarse task failure.
arXiv:2607. 00269v1 Announce Type: new Abstract: LLMs, solvers, and agent teams increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale, infeasible, conflicting, or destructive of the evidence that triggered a repair.