Hugging Face Trending Papers

Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks

Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic specifications into automata restricting an agent's actions. We argue this is the wrong product.

arXiv Machine Learning
Jun 15

Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning

arXiv:2606. 14130v1 Announce Type: new Abstract: Safe coordination problems surface in multi-agent reinforcement learning when global safety cannot be enforced by any agent unilaterally: the admissibility of one agent's action may depend on the dynamics of other agents.

By Omar Adalat, Edwin Hamel-De le Court, Francesco Belardinelli
arXiv AI
Sep 17

Autonomy in Check: Governor-Mediated Adaptive Security at the Edge

The paper proposes a split‑control architecture for adaptive security at the network edge, where an untrusted planner emits typed security intents that are vetted by a deterministic governor before being enacted. The governor checks each intent against safety, resource, temporal‑stability, and proportionality invariants, issuing signed receipts for admitted actions that are compiled into eBPF map updates. Experiments on a Raspberry Pi 5 connected to a university 5G test network show the governor can admit, reject, and bound intents at microsecond cost without disrupting protected‑flow regularity.

By Ijaz Ahmad, Ijaz Ahmad, Flavio Esposito, Erkki Harjula
arXiv Machine Learning
Jul 7

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

arXiv:2506. 07468v4 Announce Type: replace Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders patch exposed vulnerabilities.

By Mickel Liu, Liwei Jiang, Yancheng Liang, Simon Shaolei Du, Yejin Choi, Tim Althoff, Natasha Jaques
arXiv AI
Aug 5

Shielding for Higher-Order Safety

arXiv:2608. 03662v1 Announce Type: new Abstract: Safety shields are runtime enforcement mechanisms that restrict the actions of a controller to guarantee safety.

By Filip Cano, Thomas A. Henzinger, Konstantin Kueffner