SafeTune is a source‑available library that consolidates four safety‑intervention paradigms—post‑hoc weight recovery, safety‑constrained fine‑tuning, gradient‑based unlearning, and inference‑time steering—into a single, configuration‑driven workflow. It offers shared interpretability, evaluation, and deployment tools, and its modular registry allows easy addition of new methods, benchmarks, judges, models, and fine‑tuning domains. The authors demonstrate SafeTune with controlled comparisons and case studies in finance and medical deployments, showing how it characterizes safety drift, evaluates interventions on refusal‑behavior and capability metrics, and supports calibrated or layered mitigation.
By Pratinav Seth, Saisab Sadhu, Anshul Kaushal, Vinay Kumar Sankarapu
The paper introduces the Alignment Flywheel, a governance‑centric hybrid multi‑agent system (MAS) that separates decision generation from safety governance. It defines a Proposer that generates candidate trajectories, a Safety Oracle stack that evaluates safety, and an Enforcement layer that applies risk policies at runtime. A governance MAS oversees monitoring, red‑teaming, verification, and versioned release management, enabling patch‑local fixes to safety failures without retraining the Proposer. The architecture is implementation‑agnostic and is demonstrated in two scenarios: a learned spatial Oracle and a clinical GenAI proxy. The authors provide open‑source code at https://github.com/decide-ugent/Alignment-Flywheel.
By Elias Malomgr\'e, Pieter Simoens
The paper discusses the trustworthiness of agentic AI systems built on large language models, highlighting new security and operational risks such as indirect prompt injection, memory contamination, and cross‑session data leakage. It categorizes failure modes, reviews mitigation strategies—including instruction hierarchies, context isolation, and constrained tool use—and introduces the Trustworthy Agent Development Lifecycle (TADL), a six‑phase framework for specification, design, training, evaluation, deployment, and monitoring. The authors note that TADL has not yet been empirically validated but offers a structured foundation for developing more secure and accountable agentic systems, and they call for improved benchmarks and future research priorities.
By Fayeq Jeelani Syed, Rehan Ahmad, Ali Al Bataineh, Aakriti Adhikari
arXiv:2608. 14590v1 Announce Type: new Abstract: LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools.
By Pierre Dantas, Lucas Cordeiro, Ehsan Nowroozi, Tihanyi Norbert
arXiv:2606. 29887v1 Announce Type: new Abstract: In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safety policies, rather than relying on predefined risk taxonomies.
By Jiacheng Zhang, Haoyu He, Sen Zhang, Shen Wang, Xiaolei Xu, Yuhao Sun, Meng Shen, Feng Liu
HarnessRisk is a lifecycle-oriented benchmark for evaluating safety in agent harnesses that manage tools, extensions, state, permissions, and external actions. It defines six operational phases—Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery—and includes 128 sandboxed cases pairing benign user objectives with adversarial instructions. Across three harnesses, six language models, and 14 configurations, attack success rates vary from 12.6% to 80.9%, with the most vulnerable phase being Harness Configuration.
"whyItMatters":"The benchmark demonstrates that safety failures can arise in multiple harness responsibilities and that even explicit risk detection does not guarantee safe action, underscoring the need for comprehensive evaluation across model and harness configurations."
By Yajing Bai, Jinhao Duan, Jie Peng, Xianfeng Wu, Sijia Liu, Song Wang, Tianlong Chen
arXiv:2605. 12729v2 Announce Type: replace-cross Abstract: Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), including incident investigation, root-cause analysis, configuration synthesis, and limited self-healing.
By Muhammad Bilal, Jon Crowcroft, Ruizhi Wang, Xiaolong Xu, Schahram Dustdar
arXiv:2606. 02965v2 Announce Type: replace Abstract: As large language models gain tool access and are deployed as autonomous agents capable of editing records, executing transactions, and modifying infrastructure, we still evaluate them based on the sole metric of task completion.
By Victor Ojewale, Suresh Venkatasubramanian
arXiv:2608. 02683v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish complex tasks.
By Zibo Xiao, Haoyu Wang, Jun Sun
arXiv:2607. 08028v1 Announce Type: new Abstract: Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context.
By Joongho Ahn, Moonsoo Kim
arXiv:2607. 19449v1 Announce Type: cross Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely unaudited.
By Aarushi Singh
The paper introduces DGEval, a benchmark of 1,678 questions designed to assess large language models (LLMs) on the International Maritime Dangerous Goods (IMDG) Code Amendment 42‑24. It evaluates 13 models from six providers, finding that while the best model surpasses human practitioners on multiple‑choice tasks, all models perform poorly on safety‑critical areas such as stowage, segregation, and regulatory recall. The study concludes that LLMs can aid compliance tasks—especially structured Dangerous Goods List lookups with web search—but human oversight and authoritative source verification remain essential for safety‑critical deployment.
By Alexander Thomas, Hubert P. H. Shum, Darren Nellis, Manli Zhu, Phatpicha Yochum, William Bartle, Daniel Wrightson