arXiv AI By Chengxiao Dai, Zhaokun Yan, Chenjun Lei, Qiao Li, Luyan Zhang

Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems

Read the original on arXiv AI →

arXiv:2607. 20005v1 Announce Type: new Abstract: In modern IT operations (IT-Ops), the cost of an incorrect repair often exceeds the cost of no action at all.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

Distinguish or Homogenize: Last-Chance Policy Identification and Risk-Budgeted Recovery under Irreversible Resource Depletion

The paper introduces the Distinguish-or-Homogenize principle, where an agent can either spend resources to differentiate between latent fault models or alter the system state so that the remaining models share a common acceptable policy, eliminating further diagnosis. This leads to the Last-Chance Policy Identification (LCPI) framework, which evaluates correctness at the reached state rather than the initial one, and defines the Last Identifiable Margin (LIM) as the boundary between distinguishing and homogenizing. For deterministic diagnostic graphs, an Exact-LIM recursion is provided, while for noisy finite-horizon recovery the authors propose Risk-Budgeted Compatibility Planning (RBCP), which searches a compatibility-aware frontier under a hard worst-case failure constraint, demonstrating improved risk-feasible recovery in microservice and MiniGrid scenarios.

By Yibo Guo, Xiaodan Wang
arXiv AI
Jul 7

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms

arXiv:2606. 20408v3 Announce Type: replace-cross Abstract: Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized.

By Hanwool Lee, Dasol Choi, Bokyeong Kim, Haon Park, Seung Geun Kim
arXiv AI
Jun 19

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

arXiv:2606. 20408v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized.

By Hanwool Lee, Dasol Choi, Bokyeong Kim, Seung Geun Kim, Haon Park